Jason Zhu, Hejian Sang, Arup De, Rohit Jain, and Yanning Chen published a retrospective on agentic reinforcement learning training for GPT-OSS, addressing challenges including exploding KL divergence and PPO on-policy integrity issues. The authors used verl as their training framework and tested GPT-OSS-20B on tasks like ReTool and gsm8k. Their work includes an attention-sink fix for FlashAttentionV3 that supports both GPT-OSS-20B and GPT-OSS-120B, along with corrections for training-inference mismatches in mixture-of-experts log-probability calculations.
No score is assigned. Sources and their independence are shown in the citation chain below.