NTH

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

AuthorsLaurin Wagner, Mario Zusag, Bernhard Thallinger

August 3, 2026 2 min read
Watch on YouTube
The one-line take

This work makes speech recognizers choose between clean and truly verbatim transcription while improving disfluency detection, word timing, and training-data creation.

Key results

79%
German disfluency F1 with frozen tag embeddings

Zero-weight-update activation using only newly trained mode-tag embeddings.

93.8%
German zero-shot event F1 after English-only fine-tuning

Full decoder fine-tuning transfers verbatim disfluency detection without German verbatim training data.

102 ms
FluencyBank word-boundary MAE

Supervised averaged cross-attention heads outperform forced alignment and WhisperX on disfluent speech.

96.1%
Verbatimize rare-word recall

Rare-word preservation on the ICSI evaluation set after transcript-conditioned verbatim reconstruction.

60%
WER attributable to style mismatch

Estimated share of reported conversational ASR WER caused by verbatim-versus-intended reference differences.

What the paper found

This paper from nyra labs argues that verbatim versus intended transcription is an uncontrolled latent variable in modern ASR, destabilizing decoding, contaminating WER, and making word timing ambiguous. Starting from OpenAI’s Whisper-medium, the authors add coverage-aware decoder mode tags that explicitly select verbatim or intended output, supervised cross-attention alignment heads for timing, and a transcript-conditioned “verbatimize” mode that restores acoustically grounded disfluencies to an intended transcript. Training only newly added tag embeddings, with Whisper’s pretrained encoder and decoder frozen, raises German disfluency event F1 to 79%; full English-only fine-tuning transfers zero-shot to German at 93.8% event F1. On FluencyBank, supervised attention produces 102 ms mean absolute word-boundary error, outperforming forced alignment at 142 ms and WhisperX at 200 ms. Verbatimize increases rare-word recall from 6.8% to 96.1% on the ICSI rare-word set, while casing perturbation helps the model copy exact transcript content. The study also finds that up to 60% of reported WER on conversational benchmarks can reflect style mismatch rather than recognition errors. Comparisons include NVIDIA’s Canary-1B and WhisperX; GPT-4o generated intended transcript variants and Claude by Anthropic assisted some coding. The central conclusion is that pretrained ASR models often possess verbatim capability already, but explicit policy controls are needed to activate it reliably across languages and downstream timing tasks.

Original abstract

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis