NTH

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

AuthorsMiniMax, :, Aili Chen, Aonian Li, Baichuan Zhou, Bangwei Gong, Binyang Jiang, Boji Dan, Changqing Yu, Chao Wang, Cheng Ma, Cheng Zhong, Cheng Zhu, Chengjun Xiao, Chengyi Yang, Chengyu Du, Chenyang Zhang, Chi Zhang, Chuangyi Huang, Chunhao Zhang, Chunhui Du, Chunyu Zhao, Congchao Guo, Da Chen, Deming Ding, Dianjun Sun, Dongyu Zhang, Enhui Yang, Fei Yu, Guang Zheng, Guodong Zheng, Guohong Li, Haichao Zhu, Haigang Zhou, Haimo Zhang, Han Ding, Hao Zhang, Haohai Sun, Haolin Lyu, Haonan Lu, Haoyu Wang, Huajie Shi, Huiyang Li, Jiacheng Chen, Jian Zhang, Jiaqi Zhuang, Jiaren Cai, Jiaxin Pan, Jiayao Li, Jiayuan Song, Jichuan Zhang, Jie Wang, Jihao Gu, Jin Zhu, Jingwei Dong, Jingyang Li, Jingyu Zhang, Jingze Zhuang, Jinhao Tian, Jinli Liu, Jinyi Hu, Jun Tao, Jun Zhang, Junbin Ruan, Junhao Xu, Junjie Yan, Junteng Liu, Junxian He, Kang Xu, Ke Ji, Ke Yang, Kecheng Xiao, Keyu Duan, Keyu Li, Le Han, Letian Ruan, Li Yuan, Lianfei Yu, Liheng Feng, Lijie Mo, Lin Li, Lingye Bao, Lingyu Yang, Lingyua...

July 1, 2026 2 min read
Watch on YouTube
The one-line take

MiniMax-M2 is a huge but sparsely activated language model built for agents, combining verifiable agent data, specialized RL, and self-improving training workflows to push real-world task performance.

Key results

229.9B
total parameters

M2 flagship model size

9.8B
activated parameters

M2 activated per token

29.2T
pretraining tokens

M2 pretraining corpus size

192K
context window

native maximum context length

56.2
SWE-bench Pro

M2.7 agentic coding benchmark score

94.2
AIME 2026

M2.7 reasoning benchmark score

What the paper found

MiniMax’s M2 series is a Mixture-of-Experts language-model line built around “mini activations, maximum real-world intelligence,” with the flagship M2 using 229.9B total parameters but only 9.8B activated per token, a 192K context window, 256 fine-grained experts, and pretraining on 29.2T tokens. The paper’s core novelty is not just scale but an agent-native stack: large verifiable data pipelines for software engineering, AppDev, terminal work, deep search, office tasks, and reasoning; Forge, a reinforcement-learning system that unifies white-box and black-box agents with windowed-FIFO scheduling, prefix-tree merging, and speculative decoding; and M2.7’s early self-evolution loop, where the model autonomously debugs training runs and edits its own scaffold. Across benchmarks, M2.7 remains competitive with Anthropic’s Claude Opus 4.6 and Sonnet 4.6, OpenAI’s GPT-5.4, and Google DeepMind’s Gemini 3.1 Pro despite the ~10B activated-parameter footprint, reaching 56.2 on SWE-bench Pro, 76.5 on SWE-bench Multilingual, 52.7 on Multi-SWE-bench, 57.0 on Terminal-Bench 2.0, 77.8 on BrowseComp, 50.0 on GDPval-AA, 46.3 on Toolathlon, 94.2 on AIME 2026, and 89.8 on GPQA-Diamond. The strongest within-series gains come from new data and RL pipelines, including a 13-point jump on NL2Repo, a 23.2-point gain on Finance Modeling Pro, and a 66.6% medal rate on MLE Bench Lite, tying Gemini 3.1 Pro while demonstrating autonomous ML-engineering and scaffold optimization.

Original abstract

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis