NTH

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

AuthorsFrancesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones

September 4, 2026 3 min read
Watch on YouTube
The one-line take

LibriBrain100 provides over 100 hours of standardized MEG speech data and benchmarks to accelerate noninvasive brain-to-text research.

Key results

104.2
Total MEG dataset

Total hours of annotated MEG recordings across 33 subjects.

80.5
Deep single-subject component

Hours recorded from the primary subject.

32
Broad subject cohort

Additional subjects contributing approximately 40 minutes each.

0.478
Cross-subject accuracy with deep data

Balanced top-10 accuracy when training includes the deeply recorded subject.

0.329
Cross-subject accuracy without deep data

Balanced top-10 accuracy without the deeply recorded subject, a difference of about 15 percentage points.

0.739
Deep-subject supervised accuracy

Balanced top-10 accuracy from the supervised decoder on the deeply recorded subject.

What the paper found

LibriBrain100 is an open benchmark for neural speech decoding that expands MEG recordings to 104.2 hours across 33 subjects, combining 80.5 hours from one deeply recorded listener with 32 additional subjects contributing roughly 40 minutes each. Its stimuli span the full Sherlock Holmes audiobook canon, phonetic-control corpora TIMIT and MOCHA-TIMIT, and 30 semantically diverse The Moth podcasts, with 306 MEG sensors, time-locked word and phoneme annotations, standard train-validation-test splits, and Python tooling distributed through Hugging Face. Using the MEG-XL foundation model, pretrained on approximately 300 hours from 800 subjects and then fine-tuned for 50-word classification, the study reaches 0.478 balanced top-10 accuracy across the broad subject cohort when deep-subject training data is included, versus 0.329 without it—about a 15 percentage-point gain. On the deeply recorded subject, a supervised decoder reaches 0.739, compared with 0.585 for MEG-XL, showing that abundant in-domain data can outperform broad pretraining. Reducing subject-specific fine-tuning to roughly 10 minutes preserves most cross-subject performance, supporting data-efficient adaptation, although the recordings involve passive listening rather than attempted or imagined speech. The release also provides raw BIDS and minimally processed HDF5 data, a public leaderboard, and reproducible competition infrastructure; the experiments ran on NVIDIA H100 GPUs.

Original abstract

We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With $\sim$80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8$\times$ more than the next comparable dataset and roughly 80$\times$ more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected $\sim$40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis