🛰️ Daily AI Frontier
‹ back to 2026-08-09

On-Policy Delta Distillation for Multilingual Math Reasoning

Research LLMs & Foundation Models

Ranking

Overall 63
Content 60
Popularity 70

Observed public metrics from 1 member.

Merged summary

TL;DR - An arXiv study of On-Policy Distillation (OPD) for multilingual math reasoning, showing that a variant using the teacher-vs-base probability gap (OPD²) yields stronger gains in Korean and Japanese than standard OPD. It matters because it offers an RL-alternative post-training recipe that reduces cross-lingual reasoning gaps.

  • OPD² replaces the plain teacher signal with the probability delta between a post-trained teacher and its base model, and consistently beats vanilla OPD in experiments on Qwen3.
  • Gains are largest for Korean and Japanese, and the approach generally narrows the English–Korean performance gap.
  • English-only OPD can still lift Korean/Japanese math performance, but tends to shift model responses toward English.
  • Implication: multilingual distillation data is needed to preserve target-language outputs, not just to improve accuracy.

Sources (1)

On-Policy Delta Distillation for Multilingual Math Reasoning

arXiv cs.CL Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han 2026-08-06 arXiv:2608.05802
Public signals Hugging Face upvotes 32
Providers: Hugging Face · Upvotes 32 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-07 14:27:30.309481 UTC

TL;DR - An arXiv study of On-Policy Distillation (OPD) for multilingual math reasoning, showing that a variant using the teacher-vs-base probability gap (OPD²) yields stronger gains in Korean and Japanese than standard OPD. It matters because it offers an RL-alternative post-training recipe that reduces cross-lingual reasoning gaps.

  • OPD² replaces the plain teacher signal with the probability delta between a post-trained teacher and its base model, and consistently beats vanilla OPD in experiments on Qwen3.
  • Gains are largest for Korean and Japanese, and the approach generally narrows the English–Korean performance gap.
  • English-only OPD can still lift Korean/Japanese math performance, but tends to shift model responses toward English.
  • Implication: multilingual distillation data is needed to preserve target-language outputs, not just to improve accuracy.
item →