Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
TL;DR - This work treats ASR transcription style as a controllable latent variable, enabling models to switch between verbatim and intended transcripts while improving word-level timing. This reduces evaluation distortion and preserves disfluencies more reliably.
- Coverage-aware task tokens raise German disfluency F1 from 10% to 79% zero-shot after English-only training.
- English-only fine-tuning outperforms baselines on verbatim accuracy, disfluency detection, and intended-mode quality across English and German.
- Supervised cross-attention fine-tuning improves timestamps for disfluent speech beyond forced-alignment baselines.
- The proposed “verbatimize” task supports scalable creation of canonical verbatim speech transcripts.