TutorialsOrdinary
Test-Time Compute and GRPO in Practice: From PPO to Critic-Free Reinforcement Learning
Summary
A deep dive into the paradigm shift from pre-training scaling laws to test-time compute. We deconstruct the mathematical derivation of DeepSeek-R1's Group Relative Policy Optimization (GRPO), critic-free architecture advantages, emergent self-reflection in long reasoning traces, and a complete, reproducible hands-on implementation.
CategoryAI Tutorials & Practice
TierOrdinary
Published
Indexed by AIQB
SourceDEV Community
AIQB record IDintel-754ebc149ff82c144a286baf