NTH

Multiplayer Interactive World Models with Representation Autoencoders

AuthorsAnthony Hu, Václav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Amélie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Norén, James Swingos, Jan Hünermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz Böhle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick Pérez

July 15, 2026 3 min read
Watch on YouTube
The one-line take

The paper builds the first real-time multiplayer world model for Rocket League, showing that a large diffusion model can simulate coordinated multi-agent gameplay with surprisingly stable long-horizon behavior.

Key results

10K
Training data

Clean Rocket League match-hours generated by self-play bots

5B
World model size

Parameters in the flagship multiplayer flow-matching transformer

20
Inference rate

Frames per second on a single NVIDIA B200 GPU

5
Stable evaluation horizon

Minutes over which distributional quality remains steady

10.7
Latent model gFID

Generation FID at the 4-second horizon, versus 104.9 for plain pixel-space modeling

0.91
Action Recoverability Ratio

Controllability score for recovering commanded actions from generated video

What the paper found

MIRA, from teams at General Intuition, Kyutai, and Epic Games, introduces a multiplayer world model that simulates four interacting players in 2v2 Rocket League rather than treating other agents as background scenery. Trained on 10K match-hours generated by Nexto bots, its approximately 5B-parameter flow-matching transformer predicts all players’ futures jointly from synchronized views and action streams. The model operates in a representation autoencoder built on frozen DINOv3-L features, using a 10 Hz latent space, diffusion forcing for exposure to imperfect rollout context, and progressive self-distillation for fast sampling. With streaming KV caching and a 20-latent context window, MIRA generates 20 fps on a single NVIDIA B200 GPU; distributional quality remains stable through 5 minutes, despite training on short clips. The latent design is crucial: at a 4-second horizon, MIRA reaches gFID 10.7, compared with 104.9 for a matched plain pixel-space model, while its Action Recoverability Ratio reaches 0.91. Multiplayer conditioning also improves shared-world coherence, keeping cars, ball trajectories, goals, demolitions, and HUD events consistent across four viewpoints. The authors find that pretrained visual features improve generative stability even when a from-scratch codec reconstructs sharper images, and that diffusion forcing substantially outperforms teacher forcing. Remaining failures include rare uncommanded boosts, drifting clocks and scores, and incorrect behavior for underrepresented events such as a stationary ball.

Original abstract

We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis