教程 / 实战普通
Test-Time Compute and GRPO in Practice: From PPO to Critic-Free Reinforcement Learning
内容摘要
A deep dive into the paradigm shift from pre-training scaling laws to test-time compute. We deconstruct the mathematical derivation of DeepSeek-R1's Group Relative Policy Optimization (GRPO), critic-free architecture advantages, emergent self-reflection in long reasoning traces, and a complete, reproducible hands-on implementation.