Do Audio Language Models Hear and Read Distinctive Features Alike?
Ranking
Overall
71
Content
90
Popularity
27
Observed public metrics from 1 member.
Merged summary
TL;DR - A study of six audio language models finds little evidence that shared decoders represent phonological features in the same direction for spoken and written phonemes. Only voicing in two Qwen2.5-Omni models showed significant cross-modal alignment, suggesting model family matters more than scale.
- The analysis covers seven distinctive features and 15 languages from 11 language families.
- Feature directions are derived from representation offsets between minimal phoneme pairs in audio and text, then compared using cosine similarity.
- Results are evaluated against model-specific random-pairing references, which vary sevenfold across models.
- Audio representations of voicing are consistent across languages in three models, including universal pairwise agreement in two.
Sources (1)
Do Audio Language Models Hear and Read Distinctive Features Alike?
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A study of six audio language models finds little evidence that shared decoders represent phonological features in the same direction for spoken and written phonemes. Only voicing in two Qwen2.5-Omni models showed significant cross-modal alignment, suggesting model family matters more than scale.
- The analysis covers seven distinctive features and 15 languages from 11 language families.
- Feature directions are derived from representation offsets between minimal phoneme pairs in audio and text, then compared using cosine similarity.
- Results are evaluated against model-specific random-pairing references, which vary sevenfold across models.
- Audio representations of voicing are consistent across languages in three models, including universal pairwise agreement in two.