🛰️ Daily AI Frontier
‹ back to 2026-07-25

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

Research Multimodal & Generative

Merged summary

TL;DR - X³-OPD transfers reasoning from a text LLM to an audio-language model through cross-modal on-policy distillation. It improves reasoning over speech, acoustic events, and conversational cues while largely preserving existing capabilities under domain shift.

  • The audio student generates reasoning trajectories from its own acoustic perception.
  • A text teacher supplies token-level guidance using matched transcripts and verified answers.
  • Training spans speech-rendered textual reasoning, complex audio-event reasoning, and dialogue involving prosody and other paralinguistic cues.
  • Experiments on four audio benchmarks report stronger audio-grounded reasoning and chain-of-thought quality.

Sources (1)

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

arXiv cs.LG Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin 2026-07-23 arXiv:2607.21550

TL;DR - X³-OPD transfers reasoning from a text LLM to an audio-language model through cross-modal on-policy distillation. It improves reasoning over speech, acoustic events, and conversational cues while largely preserving existing capabilities under domain shift.

  • The audio student generates reasoning trajectories from its own acoustic perception.
  • A text teacher supplies token-level guidance using matched transcripts and verified answers.
  • Training spans speech-rendered textual reasoning, complex audio-event reasoning, and dialogue involving prosody and other paralinguistic cues.
  • Experiments on four audio benchmarks report stronger audio-grounded reasoning and chain-of-thought quality.
item →