X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
TL;DR - X³-OPD transfers reasoning from a text LLM to an audio-language model through cross-modal on-policy distillation. It improves reasoning over speech, acoustic events, and conversational cues while largely preserving existing capabilities under domain shift.
- The audio student generates reasoning trajectories from its own acoustic perception.
- A text teacher supplies token-level guidance using matched transcripts and verified answers.
- Training spans speech-rendered textual reasoning, complex audio-event reasoning, and dialogue involving prosody and other paralinguistic cues.
- Experiments on four audio benchmarks report stronger audio-grounded reasoning and chain-of-thought quality.