NTH

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

AuthorsNVIDIA, :, Aaron Blakeman, Aaron Thomas, Aastha Jhunjhunwala, Abhibha Gupta, Abhinav Khattar, Adam Rajfer, Adi Renduchintala, Adil Asif, Aditya Vavre, Adriana Flores Miranda, Ahmad Bilal, Aileen Zaman, Ajay Hotchandani, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Alex Gronskiy, Alex Kondratenko, Alex Steiner, Alex Ye, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi, Alice Gatti, Alisa Liu, Alok Kumar, Amar Phanishayee, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Anahita Bhiwandiwalla, Ananth Subramaniam, Andrea Santilli, Andrew Fulks, Andrew McHarg, Andrew Tao, Andrii Skliar, Anjulie Agrusa, Ankur Srivastava, Ankur Verma, Anna Shors, Anna Warno, Antoni-Joan Solergibert I Llaquet, Arham Mehta, Arkadiusz Nowaczynski, Arti Jain, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asit Mishra, Asma Kuriparambil Thekkumpate, Atefeh Sohrabizadeh, Avinash Kaur, Avinash Vem, Ayush Dattagupta, Barath Subramaniam Anandan, Bardiya Sadeghi, Ben Lanir, Benedik...

June 19, 2026 2 min read
Watch on YouTube
The one-line take

Nemotron 3 Ultra is a massive open AI model that mixes MoE, Mamba, and transformer ideas to deliver faster long-context reasoning for autonomous agents.

Key results

550B
total parameters

Nemotron 3 Ultra model scale

55B
active parameters

Parameters active per token in the MoE model

20T
pretraining tokens

Text tokens used for base pretraining

1M
context length

Extended long-context capability

5.9x
throughput gain

Higher inference throughput versus GLM-5.1-754B-A40B on 8K input / 64K output

1.46x
MTP speedup

Rollout-generation speedup from Multi-Token Prediction

What the paper found

Nemotron 3 Ultra, from NVIDIA, is a 550 billion total-parameter Mixture-of-Experts hybrid Mamba-Attention model with 55 billion active parameters per token, pretrained on 20 trillion text tokens and extended to a 1M-token context window for agentic reasoning. The paper’s core novelty is the combination of LatentMoE, Multi-Token Prediction, NVFP4 pretraining, unified RLVR, and multi-teacher on-policy distillation, which together target long-horizon autonomous workflows rather than chat-only generation. On the base model, Nemotron 3 Ultra reports 79.07 MMLU-Pro, 50.00 GPQA, 83.84 HumanEval pass@1, and 76.83 on RULER 1M, while the post-trained model reaches 56.4 on Terminal Bench 2.1, 46.7 on GDPVal, 70.7 on SWE-Bench Verified, 92.9 on TauBench Telecom, 44.4 on BrowseComp, and 92.3 on IMOAnswerBench with tools. For efficiency, NVIDIA reports up to 5.9× higher inference throughput than GLM-5.1-754B-A40B at 8K input / 64K output, with MTP boosting rollout generation by 1.46× and MTP boosting draft acceptance lengths by 3.15% to 5.82% depending on task. The final quantized release uses a 5.03 bits-per-element NVFP4 recipe and is open-sourced with base, post-trained, and quantized checkpoints plus training data and recipes on HuggingFace.

Original abstract

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD). Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE, Multi Token Prediction (MTP), NVFP4 pre-training, multi-environment RLVR, MOPD, and reasoning budget control. Nemotron 3 Ultra achieves up to ~6x higher inference throughput as compared to state-of-the-art publicly available LLMs while attaining on-par accuracy. The state-of-the-art accuracy, high inference throughput, and 1M token context length make Nemotron 3 Ultra ideal for long-running autonomous agentic tasks. We open-source the base, post-trained, and quantized checkpoints, along with the training data and recipe on HuggingFace.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis