NTH

Twin: Playing an Unknown Game with a Test-Time Digital Twin

AuthorsAlexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori

August 21, 2026 2 min read
Watch on YouTube
The one-line take

An AI agent learns the rules and goals of unfamiliar games by building and debugging its own executable digital twin.

Key results

179
ARC-AGI-3 levels cleared

Twin clears 179 of 183 levels.

93.3
Mean action-efficiency score

Twin’s mean ARC-AGI-3 score.

61.1
No-Twin ablation score

The same coding agent without the Twin harness.

87.2%
First goal hypothesis accuracy

Correct before reward on cleared levels.

2.60B
Processed tokens

Total test-time computation across 25 games.

What the paper found

Twin is a test-time world-model inference system for ARC-AGI-3, where an agent must discover both game mechanics and hidden win conditions from interaction alone. Running OpenAI Codex with GPT-5.6 Sol, Twin writes an executable Python simulator and goal predicate, validates the simulator against every observed transition, repairs mismatches through counterexamples, searches for candidate goals, and executes only plans verified inside the digital twin. On 25 ARC-AGI-3 games, it clears 179 of 183 levels and achieves a mean action-efficiency score of 93.3, compared with 61.1 for the same coding agent without the Twin harness and 7.8 for direct play. The first inferred goal is correct before reward on 87.2% of cleared levels, while 92.9% of submitted actions follow routes already tested in simulation. The main limitation is goal inference rather than dynamics: the learned twin generalizes mechanics, but hidden objectives can trigger costly exploration. This performance requires substantial test-time computation—2.60B processed tokens—showing Twin’s central tradeoff: more offline reasoning and fewer real actions.

Original abstract

We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditional approaches hand-engineer such models, one custom design per task. Each game hides its rules and goal, and our system constructs them from simulation and interaction alone. Its inductive prior over grid games is strong enough to recover the true transitions of the game and the goal on nearly all levels. Replay validation happens in a twin world model. The harness enforces that an action is not made until the program reproduces every previous observed game transition. Each mismatch between a world model prediction and the actual action result becomes a counterexample that is used to repair the world model. Twin clears 179 out of 183 levels (97.8%), and does so more efficiently than humans in 158 out of 179 levels (88.3%). The system infers the goal before any reward on 156 of the levels it clears (87.2%), and in the remaining levels automatically discovers the goal by search. The benchmark scores completion and action efficiency, between 0 and 100, against humans playing each game for the first time. Played directly, the base model scores only 7.8%; an off-the-shelf harness increases it to 61.1%, whereas our twin world model increases the same base model to 93.3%, clearing 23 out of 25 games. Building a usable world model is simpler than anticipated, whereas the harder problem is inferring the right goal.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis