🛰️ Daily AI Frontier
‹ back to 2026-09-23

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Research Efficiency & Systems

Ranking

Overall 75
Content 95
Popularity 29

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper introduces on-policy distillation (OPD) to recover long-form reasoning degraded by sub-3-bit quantization. By supervising quantized models on their own generated trajectories, OPD substantially improves math and code performance over standard teacher-forced quantization-aware distillation.

  • OPD targets quantization-amplified exposure bias, where small deviations compound during autoregressive generation and can produce repetitive loops.
  • A frozen full-precision teacher provides dense token-level guidance and task-verifier rewards on prefixes generated through the deployment-time quantized path.
  • Across four models at 2.79 and 1.88 effective bits, average BF16 performance retention rose from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval.
  • The method preserved short-form performance and outperformed continued teacher-forced distillation under matched training budgets.

Sources (1)

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

arXiv cs.LG Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng 2026-09-22 arXiv:2609.26708
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:16:43.155445 UTC

TL;DR - This paper introduces on-policy distillation (OPD) to recover long-form reasoning degraded by sub-3-bit quantization. By supervising quantized models on their own generated trajectories, OPD substantially improves math and code performance over standard teacher-forced quantization-aware distillation.

  • OPD targets quantization-amplified exposure bias, where small deviations compound during autoregressive generation and can produce repetitive loops.
  • A frozen full-precision teacher provides dense token-level guidance and task-verifier rewards on prefixes generated through the deployment-time quantized path.
  • Across four models at 2.79 and 1.88 effective bits, average BF16 performance retention rose from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval.
  • The method preserved short-form performance and outperformed continued teacher-forced distillation under matched training budgets.
item →