The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
AuthorsGilad D. Landau, Dulhan Jayalath, Oiwi Parker Jones
AffiliationsPNPL University of Oxford
Resources
The work decodes what someone means from brain signals by translating MEG activity into semantic representations before generating text.
Key results
Total MEG recording duration used for training, validation, and testing.
Sentence-level decoding score without word-level alignment.
Point improvement over the noise-control baseline.
Approximate training-data duration where performance gains begin to saturate.
Embedding-space robustness score under controlled perturbations.
What the paper found
The paper introduces Brain2Semantics2Text, a non-invasive speech-decoding system that treats meaning, rather than phonemes or individual words, as the decoding target. It processes sentence-length MEG responses with spatial attention, dilated temporal convolutions, gated linear units, and a four-head temporal Transformer, then maps the resulting representation into a pre-trained semantic embedding space. The predicted vector is converted back into text through iterative embedding inversion, avoiding word-level alignment and closed-vocabulary supervision. Among candidate spaces, the ADA embedding was selected for its strong soft reversibility, measured at 0.925, and lower sensitivity to sentence length; training combines SigLIP-style contrastive alignment with VICReg invariance, variance, and covariance regularization plus global cosine alignment. Experiments use 62.54 hours of single-subject Sherlock Holmes MEG recordings from LibriBrain. On held-out sessions, the method reaches a BERTScore of 0.830 without word-level alignment, outperforming the sentence-level BrainECHO baseline on semantic and lexical-overlap measures, although BrainECHO uses an acoustic latent space and OpenAI’s Whisper for text generation. The neural signal contributes a 6.0-point uplift in ADA cosine similarity over noise control. Scaling experiments show that retrieval performance continues improving with more data before gains begin to saturate around 55.6 hours. The results support a semantic bottleneck for extracting sentence-level meaning from noisy MEG, while also exposing limitations from corpus bias, single-subject training, imperfect embedding inversion, and the gap between semantic similarity and exact transcription.
Original abstract
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.