OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses.
ProofPaper ↗
Key points
- Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits.
- Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself.
- Quantization errors can therefore move the model into states that are absent from offline recovery data.
- On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively.
Sources (1)
- [1]OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:05 AM
We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses.
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Sep 28, 2026unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
- Sep 22, 2026vllm-project/vllm v0.30.0
- Aug 22, 2026sgl-project/sglang v0.5.18
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jun 29, 2026vllm-project/vllm v0.24.0