Kimi K3: Open Frontier Intelligence
AuthorsKimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen, Yanru Chen, Yifei Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu, Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He, Xinyi ...
Resources
Kimi K3 is an open 2.8-trillion-parameter multimodal model designed to deliver frontier-level reasoning, coding, and long-horizon agent capabilities at unprecedented scale.
Key results
Total parameters in the native multimodal Mixture-of-Experts model.
Parameters activated for each inference computation.
Maximum supported token context for long-horizon reasoning and agentic tasks.
Reported overall improvement over Kimi K2.
Coding-agent benchmark score, reported as the best result in the comparison.
Agentic web-browsing benchmark score.
What the paper found
Moonshot AI’s Kimi Team introduces Kimi K3, an open native-multimodal Mixture-of-Experts model with 2.8T total parameters, 104B activated parameters, and a 1M-token context window. Its architecture combines Kimi Delta Attention for efficient long-sequence recurrence, periodic Gated MLA global attention—building on the latent key-value ideas of DeepSeek-V2—and Attention Residuals for selective information flow across depth. Stable LatentMoE routes each token through 16 of 896 experts, while SiTU-GLU, Quantile Balancing, MoonEP, FlashKDA, and KDA Context Parallelism address numerical stability and trillion-parameter systems efficiency; together, the design claims a 2.5× scaling-efficiency improvement over Kimi K2. Post-training combines supervised fine-tuning, reinforcement learning across coding, general agents, and reasoning, three effort levels, and Multi-Teacher On-Policy Distillation, with long-horizon trajectories reaching millions of context tokens. Kimi K3 matches or exceeds leading systems on several tasks, scoring 77.8% on ProgramBench and 91.2% on BrowseComp, while native vision and Python tool use raise Math-Vision performance to 97.8%. It remains behind Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol overall, but outperforms many open and proprietary baselines, including Claude Opus 4.8, GPT-5.5, and GLM-5.2. The release of full model weights positions Kimi K3 as a major open alternative for coding, agentic workflows, reasoning, and multimodal research.
Original abstract
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.