GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
cs.CV, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: We will release our dataset, annotator, and benchmark to facilitate future research. Github Repo: https://github.com/TencentARC/GameHorizon & Project Page: https://gamehorizon-suite.github.io
Code: https://github.com/TencentARC/GameHorizon
Project page: https://gamehorizon-suite.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing
- Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- MineStudio: A Streamlined Package for Minecraft AI Agent Development
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gemini: A Family of Highly Capable Multimodal Models
- GPT-4 Technical Report
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Cradle: Empowering Foundation Agents Towards General Computer Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
- Kimi K3: Open Frontier Intelligence
- Qwen3 Technical Report
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Ovis-U1 Technical Report
- InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models