Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Merged summary
TL;DR - PINT fine-tunes a self-supervised speech encoder using parallel utterances and augmentations to produce tokens that preserve linguistic content while removing speaker and acoustic variation. This yields more consistent, efficient semantic targets for speech modeling and audio codecs.
- Aligns recordings of identical words across speakers and conditions to isolate shared linguistic content.
- Preserves frame-level temporal grounding, unlike ASR text, enabling drop-in use as audio-codec targets.
- Reduces speaker probe accuracy from 93.1% to 1.2%, a 98.7% relative reduction.
- Achieves 42% lower ABX error and 27–30% lower language-model perplexity than baselines.
Sources (1)
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
TL;DR - PINT fine-tunes a self-supervised speech encoder using parallel utterances and augmentations to produce tokens that preserve linguistic content while removing speaker and acoustic variation. This yields more consistent, efficient semantic targets for speech modeling and audio codecs.
- Aligns recordings of identical words across speakers and conditions to isolate shared linguistic content.
- Preserves frame-level temporal grounding, unlike ASR text, enabling drop-in use as audio-codec targets.
- Reduces speaker probe accuracy from 93.1% to 1.2%, a 98.7% relative reduction.
- Achieves 42% lower ABX error and 27–30% lower language-model perplexity than baselines.