ResearchResearch paperLarge Language Models · Efficiency & Inference1 source · Oct 7, 2026

OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses.

Key points

  • Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits.
  • Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself.
  • Quantization errors can therefore move the model into states that are absent from offline recovery data.
  • On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively.

Sources (1)

  • [1]OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:05 AM
    We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses.
    Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Sep 28, 2026unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
  3. Sep 22, 2026vllm-project/vllm v0.30.0
  4. Aug 22, 2026sgl-project/sglang v0.5.18
  5. Aug 10, 2026vllm-project/vllm v0.27.0
  6. Jun 29, 2026vllm-project/vllm v0.24.0

Related