DiG-bench: Discovery in Games
| Source: arXiv AI
Tags: benchmarks, scientific discovery, AI reasoning, agentic AI, rule induction, DiG-bench
DiG-bench (Discovery in Games) tests AI agents on 70 games where all rules and win conditions are unknown — every game solved by humans on first attempt yet the hardest tier defeats all frontier models, authored by researchers from DeepMind, MIT, Cambridge, and Princeton.
Details
Discovery in Games (DiG-bench) targets active rule induction — formulating novel generalizations through experimentation — rather than answering questions with known solutions. The benchmark has 70 independent games, each encoded as a short string with transformation rules that must be discovered through interaction; even the win conditions for each level are unknown. Seven difficulty tiers span from routinely solvable by frontier models to tiers that challenge the best models in agentic harnesses. All 70 games were verified solvable by at least one human on first attempt. 21 of 70 games are released publicly; the remaining 49 are held private to prevent training contamination — a deliberate design choice that preserves evaluation integrity. The author list includes Tri Dao, Jürgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas Griffiths, and Zeb Kurth-Nelson, spanning DeepMind, MIT, Cambridge, and Princeton. The benchmark addresses a gap in current AI evaluation: most benchmarks test recall or multi-step reasoning over known information. DiG-bench tests the ability to form generalizations from scratch through active experimentation — a capability closer to scientific discovery and meaningfully different from pattern retrieval.