教程 / 实战普通
Multi-Reward RL, Part 3: GDPO + CISPO + REPO-R at 27B, and the Advantage Floor That Stopped Learning
内容摘要
Qwen3.8-27B, 600 steps, 12 reward channels: CISPO, REPO-R and prompt filtering lifted holdout from −0.47 to +1.80, and an advantage floor silently zeroed 96% of the signal.