SFT vs GRPO

Опубликовано: 02 Июнь 2026
на канале: Trelis Research
5,514
169

📜Get repo access at Trelis.com/ADVANCED-fine-tuning

Tip: If you subscribe here on YouTube, click the bell to be notified of new vids

🛠 Build & Deploy Faster
Fine-tuning, Inference, Audio, Evals, and Vision Tools: https://trelis.com

💡 Need Technical or Market Assistance?
Book a Consult Here: https://forms.gle/wJXVZXwioKMktjyVA

🤝 Are You a Top Developer?
Join the Trelis team: https://trelis.com/developer-collabor...

💸 Starting a New Project/Venture?
Apply for a Trelis Grant: https://trelis.com/trelis-ai-grants/

📧 Get Trelis AI Tutorials by Email
Subscribe on Substack: https://trelis.substack.com

📸 Thumbnail Tutorial
See How It’s Made:    • Fine Tune Flux Diffusion Models with Your ...  

Video Links:
Slides: https://docs.google.com/presentation/...
DeepSeek paper: https://arxiv.org/pdf/2501.12948
Unsloth script: https://colab.research.google.com/git...
Will Brown's GRPO: https://gist.github.com/willccbb/4676...

TIMESTAMPS:
00:00 Introduction to GRPO and Study Overview
00:51 Detailed Study Methodology
03:21 Supervised Fine Tuning (SFT) Explained
04:57 Odds Ratio Preference Optimization (ORPO)
07:00 Group Relative Policy Optimization (GRPO)
10:08 Implementation and Code Walkthrough
16:16 Training Data Creation and Optimization
19:22 Analyzing and Comparing Results
20:28 Setting Up and Running GRPO
27:19 Understanding Batch Sizes and Backpropagation
27:51 Setting Up the GRPO Trainer
28:31 Exploring Reward Functions
29:12 Densifying Rewards for Better Training
31:44 Implementing GRPO Training
33:59 Running Inference and Analyzing Results
35:33 Challenges and Considerations in GRPO
44:46 Comparing GRPO with Other Techniques
48:33 Practical Recommendations for Reinforcement Learning
54:58 Conclusion and Further Resources