Researchers applied Quantization-Aware Healing (QAH) to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, producing a 4-bit model that outperforms its bfloat16 source on 7 of 9 benchmarks. QAH distills directly from the original pre-compression model rather than the recovered checkpoint, using KL divergence on logits. Gains include +7.4 on AA-LCR long-context reasoning and +5.6 on AIME 2025; it also surpasses the full-size teacher on LiveCodeBench (66.5 vs. 66.0).
No score is assigned. Sources and their independence are shown in the citation chain below.