← Back to the wire

Peer-Predictive Self-Training for Language Model Reasoning

AnnouncementResearchJul 10, 2026

Shi Feng proposes Peer-Predictive Self-Training (PST), a label-free fine-tuning framework where multiple language models improve collaboratively using cross-model aggregate responses as internal training signals. On mathematical reasoning benchmarks, PST improves exact-match accuracy by 2.2–4.3 percentage points across Gemma-2-2B, LLaMA-3.2-1B, and Qwen2.5-1.5B while reducing the generator–verifier gap by 26–40%, requiring no external supervision or teacher–student hierarchy.

Receipt № 4611 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Shi FengPerson
Canonical: https://arxiv.org/abs/2604.13356