Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Ranking
Overall
75
Content
80
Popularity
63
Observed public metrics from 1 member.
Merged summary
TL;DR - PINT fine-tunes a self-supervised speech encoder using parallel utterances and augmentations to produce tokens that preserve linguistic content while removing speaker and acoustic variation. This yields more consistent, efficient semantic targets for speech modeling and audio codecs.
- Aligns recordings of identical words across speakers and conditions to isolate shared linguistic content.
- Preserves frame-level temporal grounding, unlike ASR text, enabling drop-in use as audio-codec targets.
- Reduces speaker probe accuracy from 93.1% to 1.2%, a 98.7% relative reduction.
- Achieves 42% lower ABX error and 27–30% lower language-model perplexity than baselines.
Sources (1)
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Public signals
Hugging Face upvotes 1
TL;DR - PINT fine-tunes a self-supervised speech encoder using parallel utterances and augmentations to produce tokens that preserve linguistic content while removing speaker and acoustic variation. This yields more consistent, efficient semantic targets for speech modeling and audio codecs.
- Aligns recordings of identical words across speakers and conditions to isolate shared linguistic content.
- Preserves frame-level temporal grounding, unlike ASR text, enabling drop-in use as audio-codec targets.
- Reduces speaker probe accuracy from 93.1% to 1.2%, a 98.7% relative reduction.
- Achieves 42% lower ABX error and 27–30% lower language-model perplexity than baselines.