NTH

AI for Games in the Foundation Model Era

AuthorsMeng Luo, Yanlin Li, Hao Li, Hongzhan Lin, Pengfei Zhou, Tianjie Ju, Ran Zhang, Yeying Jin, Mong-Li Lee, Wynne Hsu

September 16, 2026 3 min read
Watch on YouTube
The one-line take

The paper maps how foundation models are transforming game playing, design, development, adaptation, and testing while emphasizing that game-specific validation remains essential.

Key results

38.3%
GameWorld Claude Sonnet 4.6 progress

Progress on 170 tasks across 34 browser games.

19.4%
GameWorld Claude Sonnet 4.6 success

Task success under keyboard-and-mouse computer use.

52.0%
GameDevBench visual-feedback pass rate

GPT-5.4 pass rate with screenshots or video, up from 41.1% without visual feedback.

88.4%
MarioGPT playable generations

Of 250 generated levels solved by A* within five attempts.

92.2%
GameGen-Verifier Acc@5

Specification-label agreement on 100 generated web games.

What the paper found

“AI for Games in the Foundation Model Era” organizes research into six roles: playing and acting, modeling players and games, designing, building and maintaining, generating and adapting at runtime, and testing and evaluating. Its central finding is that foundation models such as OpenAI’s GPT systems, Anthropic’s Claude, Google’s Gemini, Meta’s CICERO, and NVIDIA’s ACE broaden interfaces through language, vision, code, memory, and tool use, but do not eliminate game-specific controls, rules, state representations, engine bindings, or player contexts. Cross-role pipelines are emerging: gameplay trajectories train world models such as GameNGen, imagined rollouts train policies such as Dreamer 4, design plans drive executable projects in DreamGarden, and play traces guide repair in Play2Code. Quantitative evidence remains strongest for bounded, execution-grounded tasks: on 170 tasks across 34 games, GameWorld reports Claude Sonnet 4.6 at 38.3% progress and 19.4% success, while human references reach 64.1% and 55.3% for novice play. In development, GameDevBench’s visual-feedback ablation raises GPT-5.4 pass rate from 41.1% to 52.0% across 333 tasks. MarioGPT produces playable Mario levels, with 88.4% of 250 generations solved by A* within five attempts. StatePlay reduces normalized state error below 0.06 and improves mechanics-fidelity judgments by 18.6 percentage points, while GameGen-Verifier reaches 92.2% Acc@5 with up to 16.6× lower wall-clock time than its baseline. The survey concludes that artifact reuse is not capability transfer: every downstream claim needs validation under its target game, interface, engine, and player population.

Original abstract

Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt player-facing experiences at runtime, and evaluate resulting artifacts. Yet these directions have evolved largely separately, obscuring which capabilities transfer across settings and which remain tied to particular games, engines, interfaces, or player populations. We organize the literature into six roles according to the immediate use of AI output: playing and acting; modeling players and games; designing games; building and maintaining games; generating and adapting at runtime; and testing and evaluating games. For each role, we examine what structure is supplied by the game or workflow, what AI learns or produces, which capabilities and artifacts transfer across settings and roles, and what evidence supports the claims. We identify cross-role connections: trajectories train world models, learned environments provide experience for agents, design specifications drive executable implementations, and play or testing feedback guides revision. However, control schemes, rules, engine interfaces, state representations, and player contexts often remain setting-specific, so downstream claims require validation in the target setting. Evaluation is most standardized for bounded game playing and selected learned environments, while persistent state in learned worlds, repeated software revision, validated player modeling, sustained runtime adaptation, and representative automated testing remain less established. The central challenge is to reuse or transfer outputs and capabilities across roles while re-establishing evidence for effectiveness in the game-specific contexts where they are used.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis