Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
AuthorsKairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang, Linxiang Gao, Yanmohan Wang, Mingzhe Zhang, Kaiyue Wen, Kaifeng Lyu, Wenguang Chen
Resources
Puro-2B shows that a reasonably capable 2B language model can be pretrained from scratch for roughly $5,000 using consumer-grade GPUs and an open recipe.
Key results
PuRo-2B is a dense decoder-only Transformer with approximately 2B parameters.
The canonical two-phase run processes 1.4T cumulative tokens.
Accelerator-only reproduction cost for the curriculum and checkpoint-averaged run.
Canonical PuRo-2B average score across the paper’s 15-benchmark evaluation.
Estimated net speedup of blockwise FP8 after accounting for its quality penalty.
What the paper found
PuRo-2B presents an open, end-to-end recipe for pretraining a dense 2B-parameter language model on consumer NVIDIA RTX 5090 GPUs rather than expensive data-center hardware. The best run processes 1.4T tokens in FP8, using blockwise E4M3 quantization, MuonH Hyperball optimization, proxy-guided data selection, and a curriculum that orders data within each source before late constant-learning-rate continuation and six-checkpoint averaging. The FP8 design follows the blockwise scaling approach associated with DeepSeek-V3 while retaining BF16 and FP32 for sensitive operations and optimizer state. Under the paper’s accelerator-only accounting boundary, the canonical model costs $6.9K and reaches a 57.81 percent mean over 15 benchmarks, approaching Qwen2.5-1.5B and outperforming Qwen2-1.5B; a uniform-data variant reaches the Qwen2-1.5B level at about $4.4K. Ablations estimate a 1.34x quality-adjusted speedup from FP8 and a 1.19x compute-equivalent gain from MuonH, although these factors come from separate scaling experiments rather than a full factorial study. The curriculum effect persists after supervised fine-tuning: GSM8K improves from 66.89 percent to 68.66 percent in focused mathematics, from 74.10 percent to 76.12 percent in scaled mathematics, and the broad 15-task macro-average rises from 54.99 percent to 56.58 percent. The release includes data manifests, code, checkpoints, and model weights, making low-cost pretraining experiments reproducible rather than merely providing an open-weight model.
Original abstract
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \$1.5M, and reproducing SmolLM3-3B needs over \$700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models in the collection differ in token budgets and selected recipe variants. Our best model is trained at a compute cost of less than \$6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol. This cost efficiency is enabled by a combination of approaches, including hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and the data recipe. Beyond the recipe itself, we provide two additional results. First, across the Puro-2B collection, we derive a Puro Cost Scaling Law that relates training cost to average model performance; the fitted law suggests that about \$4.4K, less than \$5,090, is sufficient to reach the performance of Qwen2-1.5B. Second, as an end-to-end case study, we examine how pretraining data curricula shape downstream performance after post-training. Such controlled studies are enabled by having access to the full pretraining pipeline rather than model weights alone. We release the full training recipe for Puro-2B, including data, code, and model weights under Apache 2.0 at https://huggingface.co/collections/thu-pacman/puro-2b.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.