GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
AuthorsYiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
AffiliationsARC Lab, Tencent · GVC Lab, Great Bay University · University of Macau · National University of Singapore · Huazhong University of Science and Technology · MMLab, CUHK
Resources
GameHorizon is a large benchmark and dataset designed to test whether AI systems can understand instructions and execute gameplay tasks over short and long time horizons.
Key results
Hours of synchronized gameplay recordings.
AAA game titles represented in the dataset.
Distinct short-, medium-, and long-horizon instructions.
Models assessed across the benchmark’s model families.
Average accuracy across the three primary offline tasks.
Percentage-point improvement from adding multi-horizon instructions.
What the paper found
GameHorizon Suite introduces a unified way to evaluate gameplay across temporal scales, from immediate controls to multi-minute strategies. Its GameHorizon-Annotator uses action-aware segmentation, VLM-based semantic refinement, and dynamic programming to build a bottom-up pyramid of short-horizon operations, medium-horizon goals, and long-horizon strategies. The resulting GameHorizon-Data contains 5000 hours of synchronized video, keyboard-mouse actions, and multi-horizon instructions across 21 AAA games, including Valorant, Minecraft, Grand Theft Auto V, and Cyberpunk 2077, with 6.184036M annotated instructions. GameHorizon-Bench combines a reproducible offline multiple-choice track with stepwise online testing that resets environments after failed subtasks, allowing failures to be localized rather than reduced to one success rate. Across 47 models, including Gemini, GPT, Claude, dedicated game agents, GUI agents, and unified multimodal models, the average offline accuracy was 64.7%, with current-action prediction harder than cross-horizon consistency. Medium- and long-horizon instructions improved future-action planning by 7.0 percentage points, while top-performing GPT-6-Astra reached 80.2% overall offline accuracy and 45.0% long-horizon online success. The results show that planning and goal decomposition remain the main bottlenecks: models generally recognize observed actions more reliably than they select the next action or execute extended strategies.
Original abstract
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.