🛰️ Daily AI Frontier
‹ back to 2026-09-26

Do Audio Language Models Hear and Read Distinctive Features Alike?

arXiv cs.CL Multimodal & Generative Yuanhao Chen, Peter Chin 2026-09-24

TL;DR - A study of six audio language models finds little evidence that shared decoders represent phonological features in the same direction for spoken and written phonemes. Only voicing in two Qwen2.5-Omni models showed significant cross-modal alignment, suggesting model family matters more than scale.

  • The analysis covers seven distinctive features and 15 languages from 11 language families.
  • Feature directions are derived from representation offsets between minimal phoneme pairs in audio and text, then compared using cosine similarity.
  • Results are evaluated against model-specific random-pairing references, which vary sevenfold across models.
  • Audio representations of voicing are consistent across languages in three models, including universal pairwise agreement in two.

view merged work →