NTH

The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding

AuthorsGilad D. Landau, Dulhan Jayalath, Oiwi Parker Jones

AffiliationsPNPL University of Oxford

September 20, 2026 2 min read
Watch on YouTube
The one-line take

The work decodes what someone means from brain signals by translating MEG activity into semantic representations before generating text.

Key results

62.54
LibriBrain Sherlock Holmes recordings

Total MEG recording duration used for training, validation, and testing.

0.830
Test-set BERTScore

Sentence-level decoding score without word-level alignment.

6.0
ADA cosine signal uplift

Point improvement over the noise-control baseline.

55.6
Data-scaling saturation

Approximate training-data duration where performance gains begin to saturate.

0.925
ADA soft reversibility

Embedding-space robustness score under controlled perturbations.

What the paper found

The paper introduces Brain2Semantics2Text, a non-invasive speech-decoding system that treats meaning, rather than phonemes or individual words, as the decoding target. It processes sentence-length MEG responses with spatial attention, dilated temporal convolutions, gated linear units, and a four-head temporal Transformer, then maps the resulting representation into a pre-trained semantic embedding space. The predicted vector is converted back into text through iterative embedding inversion, avoiding word-level alignment and closed-vocabulary supervision. Among candidate spaces, the ADA embedding was selected for its strong soft reversibility, measured at 0.925, and lower sensitivity to sentence length; training combines SigLIP-style contrastive alignment with VICReg invariance, variance, and covariance regularization plus global cosine alignment. Experiments use 62.54 hours of single-subject Sherlock Holmes MEG recordings from LibriBrain. On held-out sessions, the method reaches a BERTScore of 0.830 without word-level alignment, outperforming the sentence-level BrainECHO baseline on semantic and lexical-overlap measures, although BrainECHO uses an acoustic latent space and OpenAI’s Whisper for text generation. The neural signal contributes a 6.0-point uplift in ADA cosine similarity over noise control. Scaling experiments show that retrieval performance continues improving with more data before gains begin to saturate around 55.6 hours. The results support a semantic bottleneck for extracting sentence-level meaning from noisy MEG, while also exposing limitations from corpus bias, single-subject training, imperfect embedding inversion, and the gap between semantic similarity and exact transcription.

Original abstract

Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis