X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
Ranking
Overall
69
Content
80
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - X³-OPD transfers reasoning from a text LLM to an audio-language model through cross-modal on-policy distillation. It improves reasoning over speech, acoustic events, and conversational cues while largely preserving existing capabilities under domain shift.
- The audio student generates reasoning trajectories from its own acoustic perception.
- A text teacher supplies token-level guidance using matched transcripts and verified answers.
- Training spans speech-rendered textual reasoning, complex audio-event reasoning, and dialogue involving prosody and other paralinguistic cues.
- Experiments on four audio benchmarks report stronger audio-grounded reasoning and chain-of-thought quality.
Sources (1)
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - X³-OPD transfers reasoning from a text LLM to an audio-language model through cross-modal on-policy distillation. It improves reasoning over speech, acoustic events, and conversational cues while largely preserving existing capabilities under domain shift.
- The audio student generates reasoning trajectories from its own acoustic perception.
- A text teacher supplies token-level guidance using matched transcripts and verified answers.
- Training spans speech-rendered textual reasoning, complex audio-event reasoning, and dialogue involving prosody and other paralinguistic cues.
- Experiments on four audio benchmarks report stronger audio-grounded reasoning and chain-of-thought quality.