Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

| Source: Hugging Face Blog

Tags: quantization, model compression, Multiverse Computing, GPT-OSS, MXFP4, LLM efficiency

Multiverse Computing's Quantization-Aware Healing (QAH) produces a 4-bit, 60B-parameter model that outperforms its full-precision bfloat16 counterpart on 7 of 9 benchmarks — inverting the normal accuracy-compression tradeoff for compressed LLMs.

Details

The standard pipeline for efficient LLM deployment — structural compression followed by 4-bit quantization — consistently damages the capabilities that matter most: reasoning, math, and code generation. A recovery step ('healing') is usually tacked on afterward, but nobody had rigorously studied how to do it when the model has already been structurally compressed, not just quantized.\n\nMultiverse Computing introduces Quantization-Aware Healing (QAH), a practical recipe that combines Quantization-Aware Distillation (QAD) with a careful scheduling strategy. Applied to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4 format, QAH produces a model that beats its own full-precision (bfloat16) 60B checkpoint on 7 of 9 benchmarks.\n\nThe key insight: standard QAT healing becomes unstable if training continues past its best point, while QAD — using KL-divergence against a frozen full-precision teacher instead of task loss — is more stable but needs tuning when both compression and quantization are stacked. Their recipe addresses this specific compound case.\n\nThis matters for practitioners running large model deployments: the 4-bit version is smaller, cheaper to run, and more accurate than the 16-bit version it was compressed from, challenging the assumption that compression always costs accuracy.