← Back to the wire

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

AchievementResearchAug 25, 2026

Researchers applied Quantization-Aware Healing (QAH) to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, producing a 4-bit model that outperforms its bfloat16 source on 7 of 9 benchmarks. QAH distills directly from the original pre-compression model rather than the recovered checkpoint, using KL divergence on logits. Gains include +7.4 on AA-LCR long-context reasoning and +5.6 on AIME 2025; it also surpasses the full-size teacher on LiveCodeBench (66.5 vs. 66.0).

Receipt № 15571 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

GPT-OSS 120BModel
Canonical: https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing