The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
AuthorsMiniMax, :, Aili Chen, Aonian Li, Baichuan Zhou, Bangwei Gong, Binyang Jiang, Boji Dan, Changqing Yu, Chao Wang, Cheng Ma, Cheng Zhong, Cheng Zhu, Chengjun Xiao, Chengyi Yang, Chengyu Du, Chenyang Zhang, Chi Zhang, Chuangyi Huang, Chunhao Zhang, Chunhui Du, Chunyu Zhao, Congchao Guo, Da Chen, Deming Ding, Dianjun Sun, Dongyu Zhang, Enhui Yang, Fei Yu, Guang Zheng, Guodong Zheng, Guohong Li, Haichao Zhu, Haigang Zhou, Haimo Zhang, Han Ding, Hao Zhang, Haohai Sun, Haolin Lyu, Haonan Lu, Haoyu Wang, Huajie Shi, Huiyang Li, Jiacheng Chen, Jian Zhang, Jiaqi Zhuang, Jiaren Cai, Jiaxin Pan, Jiayao Li, Jiayuan Song, Jichuan Zhang, Jie Wang, Jihao Gu, Jin Zhu, Jingwei Dong, Jingyang Li, Jingyu Zhang, Jingze Zhuang, Jinhao Tian, Jinli Liu, Jinyi Hu, Jun Tao, Jun Zhang, Junbin Ruan, Junhao Xu, Junjie Yan, Junteng Liu, Junxian He, Kang Xu, Ke Ji, Ke Yang, Kecheng Xiao, Keyu Duan, Keyu Li, Le Han, Letian Ruan, Li Yuan, Lianfei Yu, Liheng Feng, Lijie Mo, Lin Li, Lingye Bao, Lingyu Yang, Lingyua...
Resources
MiniMax-M2 is a huge but sparsely activated language model built for agents, combining verifiable agent data, specialized RL, and self-improving training workflows to push real-world task performance.
Key results
M2 flagship model size
M2 activated per token
M2 pretraining corpus size
native maximum context length
M2.7 agentic coding benchmark score
M2.7 reasoning benchmark score
What the paper found
MiniMax’s M2 series is a Mixture-of-Experts language-model line built around “mini activations, maximum real-world intelligence,” with the flagship M2 using 229.9B total parameters but only 9.8B activated per token, a 192K context window, 256 fine-grained experts, and pretraining on 29.2T tokens. The paper’s core novelty is not just scale but an agent-native stack: large verifiable data pipelines for software engineering, AppDev, terminal work, deep search, office tasks, and reasoning; Forge, a reinforcement-learning system that unifies white-box and black-box agents with windowed-FIFO scheduling, prefix-tree merging, and speculative decoding; and M2.7’s early self-evolution loop, where the model autonomously debugs training runs and edits its own scaffold. Across benchmarks, M2.7 remains competitive with Anthropic’s Claude Opus 4.6 and Sonnet 4.6, OpenAI’s GPT-5.4, and Google DeepMind’s Gemini 3.1 Pro despite the ~10B activated-parameter footprint, reaching 56.2 on SWE-bench Pro, 76.5 on SWE-bench Multilingual, 52.7 on Multi-SWE-bench, 57.0 on Terminal-Bench 2.0, 77.8 on BrowseComp, 50.0 on GDPval-AA, 46.3 on Toolathlon, 94.2 on AIME 2026, and 89.8 on GPQA-Diamond. The strongest within-series gains come from new data and RL pipelines, including a 13-point jump on NL2Repo, a 23.2-point gain on Finance Modeling Pro, and a 66.6% medal rate on MLE Bench Lite, tying Gemini 3.1 Pro while demonstrating autonomous ML-engineering and scaffold optimization.
Original abstract
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.