LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
AuthorsFrancesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
Resources
LibriBrain100 provides over 100 hours of standardized MEG speech data and benchmarks to accelerate noninvasive brain-to-text research.
Key results
Total hours of annotated MEG recordings across 33 subjects.
Hours recorded from the primary subject.
Additional subjects contributing approximately 40 minutes each.
Balanced top-10 accuracy when training includes the deeply recorded subject.
Balanced top-10 accuracy without the deeply recorded subject, a difference of about 15 percentage points.
Balanced top-10 accuracy from the supervised decoder on the deeply recorded subject.
What the paper found
LibriBrain100 is an open benchmark for neural speech decoding that expands MEG recordings to 104.2 hours across 33 subjects, combining 80.5 hours from one deeply recorded listener with 32 additional subjects contributing roughly 40 minutes each. Its stimuli span the full Sherlock Holmes audiobook canon, phonetic-control corpora TIMIT and MOCHA-TIMIT, and 30 semantically diverse The Moth podcasts, with 306 MEG sensors, time-locked word and phoneme annotations, standard train-validation-test splits, and Python tooling distributed through Hugging Face. Using the MEG-XL foundation model, pretrained on approximately 300 hours from 800 subjects and then fine-tuned for 50-word classification, the study reaches 0.478 balanced top-10 accuracy across the broad subject cohort when deep-subject training data is included, versus 0.329 without it—about a 15 percentage-point gain. On the deeply recorded subject, a supervised decoder reaches 0.739, compared with 0.585 for MEG-XL, showing that abundant in-domain data can outperform broad pretraining. Reducing subject-specific fine-tuning to roughly 10 minutes preserves most cross-subject performance, supporting data-efficient adaptation, although the recordings involve passive listening rather than attempted or imagined speech. The release also provides raw BIDS and minimally processed HDF5 data, a public leaderboard, and reproducible competition infrastructure; the experiments ran on NVIDIA H100 GPUs.
Original abstract
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With $\sim$80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8$\times$ more than the next comparable dataset and roughly 80$\times$ more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected $\sim$40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.