Twin: Playing an Unknown Game with a Test-Time Digital Twin
AuthorsAlexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori
Resources
An AI agent learns the rules and goals of unfamiliar games by building and debugging its own executable digital twin.
Key results
Twin clears 179 of 183 levels.
Twin’s mean ARC-AGI-3 score.
The same coding agent without the Twin harness.
Correct before reward on cleared levels.
Total test-time computation across 25 games.
What the paper found
Twin is a test-time world-model inference system for ARC-AGI-3, where an agent must discover both game mechanics and hidden win conditions from interaction alone. Running OpenAI Codex with GPT-5.6 Sol, Twin writes an executable Python simulator and goal predicate, validates the simulator against every observed transition, repairs mismatches through counterexamples, searches for candidate goals, and executes only plans verified inside the digital twin. On 25 ARC-AGI-3 games, it clears 179 of 183 levels and achieves a mean action-efficiency score of 93.3, compared with 61.1 for the same coding agent without the Twin harness and 7.8 for direct play. The first inferred goal is correct before reward on 87.2% of cleared levels, while 92.9% of submitted actions follow routes already tested in simulation. The main limitation is goal inference rather than dynamics: the learned twin generalizes mechanics, but hidden objectives can trigger costly exploration. This performance requires substantial test-time computation—2.60B processed tokens—showing Twin’s central tradeoff: more offline reasoning and fewer real actions.
Original abstract
We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditional approaches hand-engineer such models, one custom design per task. Each game hides its rules and goal, and our system constructs them from simulation and interaction alone. Its inductive prior over grid games is strong enough to recover the true transitions of the game and the goal on nearly all levels. Replay validation happens in a twin world model. The harness enforces that an action is not made until the program reproduces every previous observed game transition. Each mismatch between a world model prediction and the actual action result becomes a counterexample that is used to repair the world model. Twin clears 179 out of 183 levels (97.8%), and does so more efficiently than humans in 158 out of 179 levels (88.3%). The system infers the goal before any reward on 156 of the levels it clears (87.2%), and in the remaining levels automatically discovers the goal by search. The benchmark scores completion and action efficiency, between 0 and 100, against humans playing each game for the first time. Played directly, the base model scores only 7.8%; an off-the-shelf harness increases it to 61.1%, whereas our twin world model increases the same base model to 93.3%, clearing 23 out of 25 games. Building a usable world model is simpler than anticipated, whereas the harder problem is inferring the right goal.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.