HRM-Text: Efficient Pretraining Beyond Scaling
AuthorsGuan Wang, Changling Liu, Chenyu Wang, Cai Zhou, Yuhao Sun, Yifei Wu, Shuai Zhen, Luca Scimeca, Yasin Abbasi Yadkori
Resources
HRM-Text claims you can pretrain a capable language model far more cheaply by swapping Transformers for a brain-inspired recurrent design and training on instruction pairs instead of internet-scale raw text.
Key results
HRM-Text model trained from scratch
unique tokens used for pretraining
wall-clock pretraining time
estimated compute cost
benchmark score for HRM-Text 1B
benchmark score for HRM-Text 1B
What the paper found
HRM-Text, from Sapient Intelligence and MIT-affiliated authors, is an existence proof that language-model pretraining can be made far more sample-efficient by co-designing architecture and objective rather than just scaling Transformers. The paper replaces a standard decoder with a hierarchical recurrent model that separates fast execution from slow strategic state, then stabilizes deep recurrence with MagicNorm and warmup deep credit assignment. Instead of raw-text causal pretraining, it trains only on instruction-response pairs with a response-only conditional loss and PrefixLM masking, so prompt tokens are fully bidirectional while responses remain autoregressive. A 1B-parameter model trained from scratch on 40B unique tokens for 46 hours on 16 H100s, with roughly $1,472 in compute cost, reaches 60.7% on MMLU, 81.9% on ARC-C, 82.2% on DROP, 84.5% on GSM8K, and 56.2% on MATH. The authors report this uses about 100-900× fewer training tokens and 96-432× less estimated compute than comparable open models such as Llama 3.2, Qwen 3.5, Gemma 3, OLMo 3, Huginn, and Ouro, while preserving competitive benchmark performance. Ablations show that response-only task-completion training, then PrefixLM, then HRM architecture each add gains, and effective-depth analyses indicate the recurrent model maintains stronger late-layer representation change than looped Transformer baselines.
Original abstract
The current pretraining paradigm for large language models relies on massive compute and internet-scale raw text, creating a significant barrier to foundational research. In contrast, biological systems demonstrate highly sample-efficient learning through multi-timescale processing, such as the functional organization of the frontoparietal loop. Taking this as inspiration, we introduce HRM-Text, which replaces standard Transformers with a Hierarchical Recurrent Model (HRM) that decouples computation into slow-evolving strategic and fast-evolving execution layers. To stabilize this deep recurrence for language modeling, we introduce MagicNorm and warmup deep credit assignment. Furthermore, instead of standard raw-text pretraining, we train exclusively on instruction-response pairs using a task-completion objective and PrefixLM masking. Serving as an empirical existence proof of efficient pretraining, a 1B-parameter HRM-Text model trained from scratch on only 40 billion unique tokens and $1,500 budget achieves 60.7% on MMLU, 81.9% on ARC-C, 82.2% on DROP, 84.5% on GSM8K, and 56.2% on MATH. Despite utilizing roughly 100-900x fewer training tokens and 96-432x less estimated compute than standard baselines, HRM-Text performs competitively with 2-7B parameter open models. These results demonstrate that co-designing architectures and objectives can radically reduce the compute-to-performance ratio, making pretraining from scratch accessible to the broader research community.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.