DiG-bench: Discovery in Games
This benchmark addresses the lack of controlled environments for evaluating discovery and experimentation in AI, providing a new resource for the research community.
DiG-bench introduces a benchmark of 70 games designed to test AI agents' ability to discover unknown rules through experimentation. The benchmark includes seven difficulty tiers, with the hardest tier challenging state-of-the-art agentic models, while all games are solvable by humans on first attempt.
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.