DiG-bench: Discovery in Games

arXiv:2608.12593 · cs.AI, cs.LG · Submitted 2026-08-12 · Read on arXiv

Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, Jürgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas L. Griffiths, Zeb Kurth-Nelson, James C. R. Whittington

Thinking About Thinking · Independent · University of Oxford · Princeton University · King Abdullah University of Science and Technology · Swiss AI Lab · inria · Massachusetts Institute of Technology

cs.AI, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

Code: https://github.com/karpathy/autoresearch

Project page: https://digbench.ai

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: DiG-bench (Discovery in Games) is a new benchmark introduced to measure the capacity for active discovery—formulating novel generalizations through experimentation—in AI systems.

Terminology

Summary

DiG-bench (Discovery in Games) is a new benchmark introduced to measure the capacity for active discovery—formulating novel generalizations through experimentation—in AI systems. The benchmark consists of 70 independent, text-based games, each encoded as a short string with unique transformation rules that must be discovered through interaction. Both the rules and the win conditions are hidden from the player, and the games are designed to test discovery in isolation, without confounds like visual perception or prior scientific knowledge.

The paper states: "Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games)."

Six design choices distinguish DiG-bench: (1) games are purely text-based to avoid visuospatial confounds, with observations as short strings; (2) games are handcrafted, novel, and mostly kept private; (3) all games are solvable by at least one human on first attempt, though they are effortful; (4) discoveries span diverse mechanisms, so the benchmark isn't solved by a single challenge; (5) many games include a creative mode sandbox for experimentation with a less restrictive step limit; (6) difficulty is calibrated to the frontier of current models.

The benchmark contains games at seven tiers of difficulty, with tier 1 being easiest and tier 7 hardest. Each game has between 1 and 16 levels, with limited steps per level and a limited number of lives. Observations are short Unicode strings, and actions are single characters, with most games having fewer than 10 possible actions. A subset of 21 games (P-1 through P-21) is released publicly, with three games per tier; the remaining 49 games are held private for secure evaluation.

Results show that model performance varies widely. The weakest model tested, Qwen 3.6 27B, beat 1 out of 70 games, while the strongest, Opus 5, beat 50 games. Taking the best model per game, models in the basic harness collectively scored 57 out of 70 games, and 9 out of 20 in the top two tiers. Agentic harnesses (including Prime Agent, Claude Code, Codex, Kimi Code, and PRO-LONG) did not improve performance over the basic harness; for example, In the case of Kimi K3 and Gemini 3.1 Pro, where a direct comparison was available, the model in the agentic harness performed no better than the same model in the basic harness.

A key control experiment gave Gemini 3.1 Pro access to the ground-truth rules of each game in natural language. This raised its win rate from 18/70 to 69/70 games, and almost eliminated its use of creative mode. The paper concludes: These data are consistent with the idea that a primary challenge of the benchmark is finding out the rules.

Human performance confirmed that all 70 games are beatable: All 70 games were beaten by at least one human on their first attempt at that game. Humans found the games challenging, with some plays lasting over an hour, and some players reported using pen and paper. In matched levels, humans and Gemini 3.1 Pro took similar numbers of steps to beat levels (humans: 49 ± 78; Gemini: 46 ± 63 steps per level, mean ± SD; game-level paired Wilcoxon p = 0.15).

The paper also includes a comparison with ARC-AGI-3, porting five of its games into the text-based framework. Gemini 3.1 Pro cleared every level of all five ported games, and the human step-efficiency advantage seen in the original ARC-AGI-3 did not reproduce, suggesting part of ARC-AGI-3's difficulty is perceptual.

The benchmark is available at https://digbench.ai, with an API for the public games. The paper discusses related work in interactive discovery environments, scientific discovery environments, and rule induction from fixed demonstrations, positioning DiG-bench as filling the gap for a targeted, controlled benchmark of active discovery.

Improvements for AI systems

Improvements to AI systems:

  1. Add a “rule-hypothesis refinement” module that explicitly maintains a set of candidate transformation rules, updates them after each action-observation pair, and prioritizes actions that maximally reduce rule uncertainty (e.g., via information gain or counterexample generation). This directly targets the paper’s finding that rule discovery is the primary bottleneck (Gemini 3.1 Pro jumped from 18/70 to 69/70 when given ground-truth rules).

  2. Implement a “creative-mode” exploration policy for open-ended environments: when the step limit is not restrictive, the AI should deliberately perform “weird” or low-probability actions (e.g., pressing keys repeatedly, combining actions in unusual orders) to expose hidden mechanics, rather than only optimizing for immediate rewards. This addresses the observation that creative mode was underused by models that failed.

  3. Add a “failure-driven rule revision” mechanism that, upon losing a level or hitting a dead end, automatically generates and tests alternative rule hypotheses (e.g., “maybe the rule is order-dependent, not state-dependent”) before retrying. This would improve performance on the 20 games in the top two tiers, where current models collectively only solved 9/20.

  4. Integrate a “step-efficiency memory” that stores successful action sequences per level type and reuses them in similar unseen games, reducing redundant exploration. This leverages the finding that humans and Gemini 3.1 Pro used similar steps per level (49 vs. 46), suggesting that efficient planning is achievable but not yet exploited by weaker models.

  5. Add a “perceptual-invariance” layer for text-based inputs: normalize Unicode strings and abstract away surface-level formatting (e.g., symbols, spacing) to focus on structural changes, since the paper shows that ARC-AGI-3’s difficulty partly stems from perception, not reasoning—and this layer would help port that insight into DiG-bench.

What the improved AI system can do:

  • Discover hidden rules in novel, text-based environments with no prior knowledge, by actively experimenting and revising hypotheses—raising win rates from 1–50/70 (current range) toward the 69/70 ceiling achieved with ground-truth rules.

  • Solve effortful, multi-level games (1–16 levels) with limited steps and lives, by balancing exploitation of known rules with exploration of untested actions, and by recovering from failures via rule-space search.

  • Transfer discovery skills across games by reusing learned structural patterns (e.g., “rules often depend on the last character” or “actions can be chained”), enabling faster mastery of new, private games in the held-out set.

  • Match human-level step efficiency in rule-induction tasks (≈46–49 steps per level) while exceeding human consistency, since the AI can systematically enumerate hypotheses without fatigue.

  • Avoid perceptual confounds by focusing on abstract state transitions, making it robust to arbitrary text encodings—useful for real-world tasks like debugging, protocol analysis, or scientific hypothesis testing from raw logs.

Sources

Related papers