Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
AuthorsJian Hu, Huiying Li, Hao Zhang, Binfeng Xu, Yifan Zhang, Shaokun Zhang, Hemil Desai, Michael Demoret, Pavlo Molchanov, Jan Kautz, Yi Dong
Molt aims to make agentic reinforcement learning as easy to modify as ordinary PyTorch code without sacrificing distributed training performance.
Key results
Approximate Python lines in the complete RL path
Model scale tested with the same asynchronous loop
Seconds per step after reduction from 329 seconds
GB of actor GPU memory, reduced from 64.7 GB
Seconds per optimizer step, compared with 109.5 for slime
What the paper found
NVIDIA’s Molt is a PyTorch-native framework for agentic reinforcement learning designed to make algorithm changes as direct as editing ordinary Python. It combines Ray, vLLM, and NVIDIA NeMo AutoModel with FSDP2 in one asynchronous loop, while preserving three correctness guarantees: training uses the exact generated token IDs, behavior-policy log-probabilities remain attached across policy versions, and multimodal mixture-of-experts routing is replayed consistently. Agents can run either through a native environment interface or unchanged OpenAI and Anthropic SDK code, including harnesses using tools and context compaction; a loopback server captures token-exact trajectories without retokenization. Estimators such as REINFORCE++, RLOO, GRPO, DAPO, and GAE are plain functions, so new objectives or filtering stages remain localized. The framework’s RL path is approximately 8.6K lines, yet the same loop scaled from dense models to a 700B MoE at expert parallelism 256. On Qwen3-30B-A3B with DAPO-Math, speculative decoding reduced generation time from 329 seconds to 64 seconds, while optimizer offload lowered peak actor memory from 64.7 GB to 46.4 GB. In a matched asynchronous comparison against the Megatron-Core-based slime stack, Molt required 119.4 ± 2.3 seconds per optimizer step versus 109.5 ± 10.3 seconds for slime, a statistically comparable result; however, the reported benchmark was throughput-only because an upstream MoE forward mismatch blocked effective policy updates.
Original abstract
Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the cost lands on the researcher at every iteration. Molt is a PyTorch-native training framework built to keep that cost small: a codebase compact and clean enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety, so the algorithm flow can be traced and changed end to end. The agent is an ordinary program, and one asynchronous loop trains multimodal and mixture-of-experts policies while never training on a token it did not generate, consistent in tokens, policy versions, and model semantics. Leanness does not cost performance: under a matched, fully asynchronous protocol, Molt is statistically comparable to a state-of-the-art Megatron-based stack. Molt is open source and provides recipes and containers at https://github.com/NVIDIA-NeMo/labs-molt.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.