Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)

(github.com)

21 points | by popopanda 3 days ago ago

1 comments