GameHorizon Suite logo

GameHorizon Suite

Multi-Horizon Data and Evaluation in Gameplay

A large-scale data and evaluation suite that measures AAA gameplay capabilities
at different temporal horizons for diverse model families.

Yiran Wang1,*,†, Xingyilang Yin1,2,6,*, Junfu Pu1,*, Guangzhi Wang1,*, Kaifeng Li2, Mingyu Ouyang1,3, Huiqiang Sun1,4, Lingen Li1,5
Cheng Cheng1, Wangbo Yu1, Honghao Chen1, Xiaodong Cun1,2,✉, Chi-Man Pun6, Zhiguo Cao4, Ying Shan1
1ARC Lab, Tencent   2GVC Lab, Great Bay University   3NUS   4HUST   5MMLab, CUHK   6University of Macau
† Project Lead   * Equal Contribution   ✉ Corresponding Author
01 / OVERVIEW
GameHorizon Suite pipeline: Annotator, Data, and Bench
We introduce GameHorizon, a data and evaluation suite spanning multiple horizons and AAA games. It serves as a unified yardstick across a broad range of model types. It consists of three key components. GameHorizon-Annotator automatically produces a three-level pyramid of short-horizon operations, medium-horizon goals, and long-horizon strategies. GameHorizon-Data contains 5,000 hours of gameplay across 21 game titles, with temporally aligned videos, actions, and multi-horizon instructions. GameHorizon-Bench provides reproducible offline and stepwise online testing. The offline track contains thousands of standardized MCQs across three primary tasks and diagnostic variants. The online track tests order-dependent causal and order-flexible thematic tasks via the verifiable subtasks for failure localization.
Component 01

GameHorizon-Annotator

GameHorizon-Annotator is a scalable and automated pipeline. It combines action-aware segmentation, bottom-up temporal merging, and VLM annotation to annotate gameplay trajectories with the three-level instructions.

AutomatedScalableBottom-up L1→L2→L3
Component 02

GameHorizon-Data

GameHorizon-Data is the first large-scale AAA game corpus with temporally aligned video recordings, player actions, and multi-horizon instructions across diverse AAA game genres. It comprises 5,000-hour videos in 21 games.

5,000 h21 game typesAAA-focused dataset
Component 03

GameHorizon-Bench

GameHorizon-Bench consists of reproducible offline evaluation, stepwise online testing, as well as fine-grained failure diagnosis across diverse model families. We test 47 models in VLMs, UMMs, GUI, coding, and game agents.

5,000 offline MCQs20 online tasks, 62 subtasks
02 / GAMEHORIZON-ANNOTATOR
Workflow of GameHorizon-Annotator: action-aware segmentation and bottom-up temporal merging into multi-horizon instructions
Workflow of GameHorizon-Annotator. Videos and actions are first processed by action-aware segmentation to produce short-horizon clips, with key actions determining their temporal boundaries. A VLM annotates each clip with an L1 operation. Lower-level clips with their instructions are progressively merged into medium- and long-horizon clips based on action continuity and semantic coherence, from which the VLM can derive the L2 goals and L3 strategies in the bottom-up procedure.
GameHorizon Suite pyramid of primitive actions, short-horizon operations, medium-horizon goals, and long-horizon strategies across diverse games
Examples of the annotated instructions. We construct a pyramid of primitive actions, short-horizon operations, medium-horizon goals, and long-horizon strategies.
03 / GAMEHORIZON-DATA
Black Myth: Wukong
Elden Ring
Cyberpunk 2077
Red Dead Redemption 2
GTA V
The Witcher 3
Assassin's Creed
Elden Ring: Nightreign
Apex Legends
PUBG
Palworld
Delta Force
Escape from Tarkov
Watch Dogs
Genshin Impact
Valorant
Minecraft
Neverness to Everness
Wuthering Waves
Roco Kingdom: World

21 AAA titles, 5,000 hours of gameplay. GameHorizon-Data is the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. The game trajectories are collected by 100 experienced human players. It spans five diverse game genres, such as the open-world, action role-playing, competitive shooter, sandbox survival, and creature-collecting adventure. It provides a reliable data foundation for evaluating gameplay abilities across models.

Game Title Recording Duration and Share Annotation Counts Avg. Instruction Span (s)
Total (h)Share (%)Valid Rate (%)Valid (h) Actions (104)L1L2L3 L1L2L3
Valorant617.512.3586.4533.85,811705,38023,1135,7782.7283.1332.5
Minecraft589.811.8084.5498.35,087448,16621,7155,4294.0082.6330.4
Grand Theft Auto V566.111.3286.5489.93,585382,30221,6445,4114.6181.5325.9
Palworld438.78.7783.5366.53,741329,61515,9713,9934.0082.6330.4
Delta Force337.06.7485.1286.93,123379,09112,4223,1052.7283.1332.5
Roco Kingdom: World329.26.5891.9302.42,677786,44613,0543,2631.3883.4333.6
Red Dead Redemption 2304.46.0985.9261.61,992202,42811,5532,8884.6581.5326.1
Cyberpunk 2077301.26.0287.9264.72,574192,22211,6462,9124.9681.8327.3
Genshin Impact232.64.6590.4210.41,862547,1349,0822,2701.3883.4333.6
PUBG: Battlegrounds197.43.9584.6167.01,818220,7267,2331,8082.7283.1332.5
Elden Ring Nightreign168.93.3890.0152.01,348250,8776,6301,6582.1882.6330.2
The Witcher 3: Wild Hunt158.73.1789.4141.81,323287,4336,2901,5731.7881.2324.6
Elden Ring135.92.7292.6125.91,116207,7065,4891,3722.1882.6330.2
Escape from Tarkov119.82.4081.797.91,065129,3204,2371,0592.7283.1332.5
Apex Legends116.52.3385.399.41,082131,3814,3051,0762.7283.1332.5
Wuthering Waves108.42.1789.196.5819294,0884,0771,0191.1885.2341.0
Neverness to Everness101.32.0389.390.4800235,1453,9039761.3883.4333.6
Assassin's Creed51.11.0287.044.533834,4011,9634914.6581.5326.1
Black Myth: Wukong51.11.0291.346.641476,9712,0345092.1882.6330.2
Watch Dogs 238.50.7785.032.724925,3241,4453614.6581.5326.1
Honor of Kings: World35.90.7287.331.327781,4291,3523381.3883.4333.6
Total / Avg.5,000100.0086.84,34141,1035,947,588189,15847,2902.6382.6330.4
Statistics of GameHorizon-Data across 21 game titles. We report the duration and share of recordings, counts of actions and instructions, and the average temporal span of each distinct instruction. Games are sorted by duration. Valid rate denotes the fraction retained for instruction annotation after filtering. Actions are in 104 unit.
Dataset Data Scope Annotation Properties Notes
Large-ScaleAAA-Focused Direct Human ActionsInstruction AnnotationsDense InstructionsMulti-Horizon
Single-Game Datasets
MineRL (Guss et al., 2019)Minecraft;
human actions;
no textual instructions.
MineDojo (Fan et al., 2022)Web Minecraft videos;
no actions or instructions.
VPT (Baker et al., 2022)Web Minecraft videos;
IDM actions;
no instructions.
STEVE-1 (Lifshitz et al., 2023)Minecraft;
only 10K short-term text instructions.
PLAICraft (He et al., 2025)Minecraft;
human actions;
no textual instructions.
WildWorld (Li et al., 2026)Monster Hunter Wilds;
AI actions;
no instructions.
EgoCS-400K (Guo et al., 2026)CS;
replay-derived action labels;
no instructions.
Multi-Game Datasets
NitroGen (Magne et al., 2026)Gamepad-overlay actions without text instructions.
Open-P2P (Yue et al., 2026)Majority non-AAA games;
sparse text instructions.
D2E (Choi et al., 2026)Mostly IDM-inferred actions;
no text instructions.
Gaming-500-hours (Markov AI, 2026)Majority non-AAA games;
no textual instructions.
GameHorizon-Data (Ours)AAA;
human actions;
dense multi-horizon instructions.
Comparison with existing datasets. Large-scale denotes at least 3,000 hours of gameplay. AAA-focused indicates that more than half of the included game titles are AAA games. The direct human actions refer to cases where the majority of the action labels are recorded directly from human players rather than inferred or generated.
04 / GAMEHORIZON-BENCH
The offline track of GameHorizon-Bench: single-horizon action, multi-horizon decomposition, and cross-horizon consistency tasks
The offline track of GameHorizon-Bench. It comprises three primary tasks and ten variant tasks. Given sampled frames and the short-horizon instruction L1, single-horizon action T1 requires models to determine the correct action sequence. Multi-horizon decomposition T2 evaluates whether models can decompose a medium-horizon goal L2 into the correct sequence of short-horizon operations L1. Cross-horizon consistency T3 assesses the overall consistency across frames, L1–L3 instructions, and actions. Variant tasks T* provide additional diagnostics and insights into different gameplay abilities. Some text are abridged due to space constraints.
Tier 1
#ModelT1T2T3Overall
1GPT-6-Astra69.479.691.580.2
2Gemini 3.8 Flash65.981.284.877.3
3Gemini 3.7 Flash66.280.183.976.7
4Gemini 3.6 Flash62.379.184.675.3
5GPT-5.6 Sol65.377.581.574.8
6Kimi-K364.479.379.774.5
7Gemini 3.5 Flash64.977.480.874.4
8GPT-5.564.678.976.273.2
9Gemini 3.1 Pro64.575.878.272.8
10Doubao-Seed-2.1-Turbo63.179.475.572.7
11Doubao-Seed-2.1-Pro62.976.776.672.1
Tier 2
#ModelT1T2T3Overall
12Doubao-Seed-2.0-Pro59.274.980.071.4
13Claude Fable 562.972.778.171.2
14GPT-5.6 Terra62.175.674.470.7
15Qwen3.7-Plus58.875.971.568.7
16Qwen3.8-27B59.977.468.168.5
17GPT-5.6 Luna61.670.771.167.8
17Kimi-K2.661.273.269.067.8
19Doubao-Seed-2.0-Lite55.269.076.366.8
20Gemma 4 31B-IT53.868.676.666.3
21MiniMax-M357.868.570.665.6
22Qwen3-VL-235B-A22B-Thinking58.271.665.965.2
Tier 3
#ModelT1T2T3Overall
23Claude Opus 4.859.162.173.164.8
24Claude Sonnet 554.457.782.064.7
25Step-3.7-Flash57.571.363.464.1
26Qwen3-VL-235B-A22B-Instruct55.162.173.163.4
27GLM-5V-Turbo53.465.470.863.2
28GPT-5.258.958.970.662.8
29Qwen3.5-397B-A17B53.360.173.562.3
30BAGEL-7B-MoT51.757.375.161.4
31Qwen2.5-VL-32B-Instruct54.858.069.760.8
32Step3-VL-10B55.464.761.760.6
33SenseNova-U1-8B-MoT48.259.267.558.3
Tier 4
#ModelT1T2T3Overall
34Qwen3.6-35B-A3B53.352.966.357.5
35Qwen3-Omni-30B-A3B-Instruct52.555.663.757.3
36GELab-Zero-4B-Preview51.846.872.757.1
37Qwen2.5-VL-7B-Instruct52.054.363.756.7
38InternVL3.5-8B51.850.265.555.8
39GLM-4.1V-9B-Thinking54.954.855.355.0
40GPT-4o55.746.361.354.4
41UI-TARS-1.5-7B49.147.162.953.0
42Ovis-U1-3B44.739.659.247.8
43InternVL3.5-2B45.140.354.746.0
44InternVL-U-4B46.137.949.844.6
Average over all modelsT157.3T265.1T371.6Overall64.7
Offline results of the primary tasks T1–T3. The table reports results for general-purpose VLMs, UMMs, GUI agents, and coding agents with question-answering abilities. Accuracies are reported as percentages. Overall denotes the mean accuracy across three tasks. Models are ordered from higher to lower performance and divided into four tiers. The best and second-best metrics are marked in bold and underlined. The final row reports the average accuracy for all models in table above.
The online track of GameHorizon-Bench: long-horizon causal and thematic tasks with verifiable short-horizon subtasks
The online track of GameHorizon-Bench. It contains long-horizon causal and thematic tasks, each comprising multiple verifiable short-horizon subtasks. The causal tasks can only be completed in a prescribed order because of dependencies between successive subtasks. The thematic tasks contain subtasks that share a common theme but can be performed in any order. The figure shows results from Gemini 3.6 Flash, with arrows indicating its actual execution order. After a subtask fails, the environment is reset to the corresponding success state so that evaluation can continue. A long-horizon task is passed only when all the constituent subtasks succeed.
Tier 1
Online RankModelOffline RankShort-Horizon SubtasksLong-Horizon TasksCausalThematic
1GPT-6-Astra166.1(41/62)45.0(9/20)40.0(4/10)50.0(5/10)
2Gemini 3.6 Flash456.5(35/62)30.0(6/20)30.0(3/10)30.0(3/10)
3Kimi-K3646.8(29/62)10.0(2/20)20.0(2/10)0.0(0/10)
Tier 2
Online RankModelOffline RankShort-Horizon SubtasksLong-Horizon TasksCausalThematic
4GPT-5.6 Terra1437.1(23/62)10.0(2/20)20.0(2/10)0.0(0/10)
5GPT-5.6 Luna1733.9(21/62)10.0(2/20)10.0(1/10)10.0(1/10)
6MiniMax-M32129.0(18/62)10.0(2/20)10.0(1/10)10.0(1/10)
Tier 3
Online RankModelOffline RankShort-Horizon SubtasksLong-Horizon TasksCausalThematic
7GLM-5V-Turbo2727.4(17/62)5.0(1/20)10.0(1/10)0.0(0/10)
8Qwen3.5-397B-A17B2919.4(12/62)5.0(1/20)10.0(1/10)0.0(0/10)
9Step3-VL-10B3214.5(9/62)5.0(1/20)10.0(1/10)0.0(0/10)
Tier 4
Online RankModelOffline RankShort-Horizon SubtasksLong-Horizon TasksCausalThematic
10Qwen3.6-35B-A3B3411.3(7/62)5.0(1/20)10.0(1/10)0.0(0/10)
11InternVL3.5-8B383.2(2/62)0.0(0/20)0.0(0/10)0.0(0/10)
12UI-TARS-1.5-7B411.6(1/62)0.0(0/20)0.0(0/10)0.0(0/10)
Online results of GameHorizon-Bench. Due to the evaluation costs, we test 12 models in our online track, with three models per offline tier. Each entry reports the success rate as a percentage, followed by the number of passed tasks out of the total in parentheses. The offline and online rankings show a clear positive association.
BenchmarkBenchmark ScopeAnnotation PropertiesEvaluation SettingsNotes
Model DiversityAAA-FocusedDirect Human ActionsMulti-HorizonReproducible OfflineStepwise Online
Single-Game Benchmarks
MineDojo (Fan et al., 2022)Online success rates in Minecraft; the game agents only.
MCU (Zheng et al., 2025)Minecraft; VLM-based online scores; game agents only.
MineExplorer (Ju et al., 2026)Online milestone rates in Minecraft; the VLMs only.
Multi-Game Benchmarks
BALROG (Paglieri et al., 2025)Non-AAA; the online progress scores; LLMs and VLMs.
VideoGameBench (Zhang et al., 2025)Retro and classic games; online scores and VLMs only.
Orak (Park et al., 2026)Majority non-AAA; the online scores; LLMs and VLMs.
GameWorld (Ouyang et al., 2026)Browser mini-games; online scores; VLM agents.
GameVerse (Zhang et al., 2026)Mostly non-AAA; online milestone scores; VLM agents.
GameHorizon-Bench (Ours)AAA; reproducible offline; stepwise online; diverse models.
Comparison with existing benchmarks. Model diversity denotes evaluation across at least three model paradigms, e.g., VLMs, UMMs, GUI, coding, and game agents.
05 / CITATION
@article{GameHorizonSuite2026,
  title   = {GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay},
  author  = {Yiran Wang and Xingyilang Yin and Junfu Pu and Guangzhi Wang and Kaifeng Li and Mingyu Ouyang and Huiqiang Sun and Lingen Li and Cheng Cheng and Wangbo Yu and Honghao Chen and Xiaodong Cun and Chi-Man Pun and Zhiguo Cao and Ying Shan},
  year    = {2026},
  journal = {arXiv preprint arXiv:2609.25001},
  url     = {https://arxiv.org/abs/2609.25001}
}