Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
AuthorsLaurin Wagner, Mario Zusag, Bernhard Thallinger
Resources
This work makes speech recognizers choose between clean and truly verbatim transcription while improving disfluency detection, word timing, and training-data creation.
Key results
Zero-weight-update activation using only newly trained mode-tag embeddings.
Full decoder fine-tuning transfers verbatim disfluency detection without German verbatim training data.
Supervised averaged cross-attention heads outperform forced alignment and WhisperX on disfluent speech.
Rare-word preservation on the ICSI evaluation set after transcript-conditioned verbatim reconstruction.
Estimated share of reported conversational ASR WER caused by verbatim-versus-intended reference differences.
What the paper found
This paper from nyra labs argues that verbatim versus intended transcription is an uncontrolled latent variable in modern ASR, destabilizing decoding, contaminating WER, and making word timing ambiguous. Starting from OpenAI’s Whisper-medium, the authors add coverage-aware decoder mode tags that explicitly select verbatim or intended output, supervised cross-attention alignment heads for timing, and a transcript-conditioned “verbatimize” mode that restores acoustically grounded disfluencies to an intended transcript. Training only newly added tag embeddings, with Whisper’s pretrained encoder and decoder frozen, raises German disfluency event F1 to 79%; full English-only fine-tuning transfers zero-shot to German at 93.8% event F1. On FluencyBank, supervised attention produces 102 ms mean absolute word-boundary error, outperforming forced alignment at 142 ms and WhisperX at 200 ms. Verbatimize increases rare-word recall from 6.8% to 96.1% on the ICSI rare-word set, while casing perturbation helps the model copy exact transcript content. The study also finds that up to 60% of reported WER on conversational benchmarks can reflect style mismatch rather than recognition errors. Comparisons include NVIDIA’s Canary-1B and WhisperX; GPT-4o generated intended transcript variants and Claude by Anthropic assisted some coding. The central conclusion is that pretrained ASR models often possess verbatim capability already, but explicit policy controls are needed to activate it reliably across languages and downstream timing tasks.
Original abstract
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.