[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 · 27:34
Josh McGrath, an OpenAI post-training researcher, states that real post-training innovation lies in data quality and signal trust, not optimization methods, with RLVR and token efficiency central. He moved from pre-training (3% compute gains) to post-training (40% behavior change), describes the infrastructure chaos of RL runs, and cites GRPO from DeepSeek Math as underappreciated for providing verifiable reward signals. GPT-5 to GPT-5.1 bumped evals while slashing tokens, emphasizing token efficiency over wall-clock time. The shopping model features interruptibility and chain-of-thought transparency, and personality toggles (Anton vs Clippy) are a key differentiator. For long context, he argues agents with graph walks may be more important than 10M-token windows. He concludes that the education system fails to produce people skilled in both distributed systems and ML research, a critical combination as bottlenecks shift.