← Back to the wire

HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning

AnnouncementResearchJul 10, 2026

Weiqi Wang proposes HeaPA, a method combining heap sampling and on-policy query augmentation to improve RLVR training efficiency for LLMs. The approach maintains an evolving prompt pool, tracks capability frontiers via heap-based boundary sampling, and grows the pool through on-policy augmentation with asynchronous validation. Across seven benchmarks, HeaPA consistently improves accuracy and reaches target performance with fewer computations at comparable wall-clock time, with more pronounced gains at mid-to-large model scales.

Receipt № 4651 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Weiqi WangPerson
Canonical: https://arxiv.org/abs/2601.22448