NTH

HRM-Text: Efficient Pretraining Beyond Scaling

AuthorsGuan Wang, Changling Liu, Chenyu Wang, Cai Zhou, Yuhao Sun, Yifei Wu, Shuai Zhen, Luca Scimeca, Yasin Abbasi Yadkori

June 12, 2026 2 min read
Watch on YouTube
The one-line take

HRM-Text claims you can pretrain a capable language model far more cheaply by swapping Transformers for a brain-inspired recurrent design and training on instruction pairs instead of internet-scale raw text.

Key results

1B
model_size

HRM-Text model trained from scratch

40B
training_tokens

unique tokens used for pretraining

46
training_time_hours

wall-clock pretraining time

1472
training_cost_usd

estimated compute cost

60.7%
MMLU

benchmark score for HRM-Text 1B

81.9%
ARC-C

benchmark score for HRM-Text 1B

What the paper found

HRM-Text, from Sapient Intelligence and MIT-affiliated authors, is an existence proof that language-model pretraining can be made far more sample-efficient by co-designing architecture and objective rather than just scaling Transformers. The paper replaces a standard decoder with a hierarchical recurrent model that separates fast execution from slow strategic state, then stabilizes deep recurrence with MagicNorm and warmup deep credit assignment. Instead of raw-text causal pretraining, it trains only on instruction-response pairs with a response-only conditional loss and PrefixLM masking, so prompt tokens are fully bidirectional while responses remain autoregressive. A 1B-parameter model trained from scratch on 40B unique tokens for 46 hours on 16 H100s, with roughly $1,472 in compute cost, reaches 60.7% on MMLU, 81.9% on ARC-C, 82.2% on DROP, 84.5% on GSM8K, and 56.2% on MATH. The authors report this uses about 100-900× fewer training tokens and 96-432× less estimated compute than comparable open models such as Llama 3.2, Qwen 3.5, Gemma 3, OLMo 3, Huginn, and Ouro, while preserving competitive benchmark performance. Ablations show that response-only task-completion training, then PrefixLM, then HRM architecture each add gains, and effective-depth analyses indicate the recurrent model maintains stronger late-layer representation change than looped Transformer baselines.

Original abstract

The current pretraining paradigm for large language models relies on massive compute and internet-scale raw text, creating a significant barrier to foundational research. In contrast, biological systems demonstrate highly sample-efficient learning through multi-timescale processing, such as the functional organization of the frontoparietal loop. Taking this as inspiration, we introduce HRM-Text, which replaces standard Transformers with a Hierarchical Recurrent Model (HRM) that decouples computation into slow-evolving strategic and fast-evolving execution layers. To stabilize this deep recurrence for language modeling, we introduce MagicNorm and warmup deep credit assignment. Furthermore, instead of standard raw-text pretraining, we train exclusively on instruction-response pairs using a task-completion objective and PrefixLM masking. Serving as an empirical existence proof of efficient pretraining, a 1B-parameter HRM-Text model trained from scratch on only 40 billion unique tokens and $1,500 budget achieves 60.7% on MMLU, 81.9% on ARC-C, 82.2% on DROP, 84.5% on GSM8K, and 56.2% on MATH. Despite utilizing roughly 100-900x fewer training tokens and 96-432x less estimated compute than standard baselines, HRM-Text performs competitively with 2-7B parameter open models. These results demonstrate that co-designing architectures and objectives can radically reduce the compute-to-performance ratio, making pretraining from scratch accessible to the broader research community.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis