On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
AuthorsMind Lab, :, Song Cao, Vic Cao, Kaijie Chen, Bunny Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Mutian Hong, Hailee Hou, Peixuan Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Autumn Jin, Fancy Kong, Kyrie Lei, Alexy Li, Dawn Li, Ray Li, Theo Li, Wenhao Li, Jiayi Lin, Domini Liu, Heshan Liu, Kairus Liu, Logan Liu, Maeve Luo, Runism Lv, Pony Ma, Verity Niu, Anson Qiu, Vincent Wang, Maxwell Yao, Regis Ye, Wenlin Ye, Yanying Ye, Josh Ying, Danney Zeng, Salmon Zhan, Anya Zhang, Ruijia Zhang, Shiyang Zhang, Sueky Zhang, Ya Zhang, Wei Zhao, Ada Zhou, Sizer Zhou, Xinyue Zhu, Murphy Zhuang
Resources
This paper argues that small adapters on top of foundation models could become persistent personal models at massive scale, enabling individualized behavior without retraining huge models.
Key results
The Kimi K2 LoRA RL case adapts a sparse MoE model with 1.04T total parameters and 32.6B activated parameters.
Trillion-scale LoRA RL is reported to reduce compute and communication to about 10% of conventional full-parameter RL.
The rank-sweep study uses 216 PPO runs across ranks, batch sizes, and seeds to characterize low-rank stability.
OLoRA-tail improves average benchmark accuracy from 56.3% to 58.3% on DeepSeek-R1-Distill-Qwen-1.5B.
At rank 1, OLoRA-tail reaches 35.5% average pass rate versus 24.0% for standard LoRA.
In the model-count experiment, accuracy rises from 0.3644 at k=1 to 0.4867 at k=198 under majority vote across distinct LoRA variants.
What the paper found
This Mind Lab paper reframes parameter-efficient fine-tuning, especially LoRA, as a mechanism for persistent local state on top of shared foundation models, aiming at “million personal models of trillion parameters.” Its core claim is that PEFT becomes transformative only when three scaling axes reinforce each other: Scale Up, where stronger shared priors make small adapters more powerful; Scale Down, where the adaptive state becomes small and stable enough to learn reliably; and Scale Out, where many persistent adapted instances can coexist. On Scale Up, the authors show trillion-scale LoRA reinforcement learning on a 1.04T-parameter sparse MoE model, Kimi K2, using GRPO-style on-policy optimization with hybrid parallelism, reducing compute and communication to about 10% of conventional full-parameter RL while preserving stable reward and task-success curves. They also identify training–inference mismatch and routing divergence in MoE systems as new failure modes that require router replay and provenance-aware serving. On Scale Down, a Qwen3-8B PPO sweep over 216 runs finds that ranks 16 and 32 are the most reliable middle regime, while ranks 1–4 retain strong best-case performance but suffer seed sensitivity; an RL-native initialization, OLoRA-tail, stabilizes rank-1 training and improves DeepSeek-R1-Distill-Qwen-1.5B benchmark accuracy from 56.3% to 58.3% average, with a larger win on Qwen3-30B-A3B-Instruct, where it reaches 35.5% versus 24.0% for standard LoRA. On Scale Out, the paper introduces LoRA-as-memory and Context Learning, reports a bounded memory capacity law on DishNameBenchmark, and shows per-user LoRA adapters improve social simulation in OASIS and majority-vote accuracy on AIME24 from 0.3644 at one model to 0.4867 at 198 models. The infrastructure layer, MinT from Mind Lab, turns adapter identity, revisioning, and bounded residency into a deployable lifecycle for persistent personal models.
Original abstract
Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We organize the problem around three scaling axes: Scale Up, where stronger shared priors make small local updates more useful; Scale Down, where we study how small adapters can be while remaining reliable; and Scale Out, where many persistent adapted instances coexist. MinT provides one infrastructure example for managing adapter identity, revision, provenance, evaluation, and serving residency. Together, the results suggest that PEFT can be a compact substrate for persistent personal models rather than only a budget substitute for full fine-tuning.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.