Do Audio Language Models Hear and Read Distinctive Features Alike?
TL;DR - A study of six audio language models finds little evidence that shared decoders represent phonological features in the same direction for spoken and written phonemes. Only voicing in two Qwen2.5-Omni models showed significant cross-modal alignment, suggesting model family matters more than scale.
- The analysis covers seven distinctive features and 15 languages from 11 language families.
- Feature directions are derived from representation offsets between minimal phoneme pairs in audio and text, then compared using cosine similarity.
- Results are evaluated against model-specific random-pairing references, which vary sevenfold across models.
- Audio representations of voicing are consistent across languages in three models, including universal pairwise agreement in two.