Shi Feng proposes Peer-Predictive Self-Training (PST), a label-free fine-tuning framework where multiple language models improve collaboratively using cross-model aggregate responses as internal training signals. On mathematical reasoning benchmarks, PST improves exact-match accuracy by 2.2–4.3 percentage points across Gemma-2-2B, LLaMA-3.2-1B, and Qwen2.5-1.5B while reducing the generator–verifier gap by 26–40%, requiring no external supervision or teacher–student hierarchy.
No score is assigned. Sources and their independence are shown in the citation chain below.