← Back to the wire

vLLM V0 to V1: Correctness Before Corrections in RL

AchievementResearchMay 7, 2026

ServiceNow-AI researchers led by Rafael Pardinas and Ehsan Kamalloo migrated PipelineRL from vLLM V0 to vLLM V1, achieving parity with their reference run after fixing four backend issues. The team corrected processed rollout logprobs, V1-specific runtime defaults, the inflight weight-update path, and the fp32 lm_head final projection. They prioritized fixing backend correctness before modifying the RL objective, ensuring train-inference logprob consistency.

Receipt № 4981 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

ServiceNow-AICompanyvLLM V0ModelvLLM V1ModelPipelineRLModelRafael PardinasPersonEhsan KamallooPersonHugging FaceCompany
Canonical: https://huggingface.co/blog/ServiceNow-AI/correctness-before-corrections