NTH

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

AuthorsYiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan

AffiliationsARC Lab, Tencent · GVC Lab, Great Bay University · University of Macau · National University of Singapore · Huazhong University of Science and Technology · MMLab, CUHK

September 25, 2026 2 min read
Watch on YouTube
The one-line take

GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.

Key results

5000
GameHorizon-Data duration

Hours of synchronized gameplay recordings.

21
Game coverage

AAA game titles represented in the dataset.

6.184036M
Annotated instructions

Distinct short-, medium-, and long-horizon instructions.

47
Evaluated models

Models assessed across the benchmark’s model families.

64.7%
Mean offline accuracy

Average accuracy across the three primary offline tasks.

7.0
Future-action planning gain

Percentage-point improvement from adding multi-horizon instructions.

What the paper found

GameHorizon Suite introduces a unified way to evaluate gameplay across temporal scales, from immediate controls to multi-minute strategies. Its GameHorizon-Annotator uses action-aware segmentation, VLM-based semantic refinement, and dynamic programming to build a bottom-up pyramid of short-horizon operations, medium-horizon goals, and long-horizon strategies. The resulting GameHorizon-Data contains 5000 hours of synchronized video, keyboard-mouse actions, and multi-horizon instructions across 21 AAA games, including Valorant, Minecraft, Grand Theft Auto V, and Cyberpunk 2077, with 6.184036M annotated instructions. GameHorizon-Bench combines a reproducible offline multiple-choice track with stepwise online testing that resets environments after failed subtasks, allowing failures to be localized rather than reduced to one success rate. Across 47 models, including Gemini, GPT, Claude, dedicated game agents, GUI agents, and unified multimodal models, the average offline accuracy was 64.7%, with current-action prediction harder than cross-horizon consistency. Medium- and long-horizon instructions improved future-action planning by 7.0 percentage points, while top-performing GPT-6-Astra reached 80.2% overall offline accuracy and 45.0% long-horizon online success. The results show that planning and goal decomposition remain the main bottlenecks: models generally recognize observed actions more reliably than they select the next action or execute extended strategies.

Original abstract

Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis