AISCJul 6

ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization

arXiv:2607.051856.5Has Code
Predicted impact top 79% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For AI researchers studying compositional generalization, this benchmark provides a structured testbed with explicit compositional strategies, addressing the lack of complex, non-linguistic benchmarks in the field.

ClassicLogic introduces a benchmark of four classic logic puzzles (Sudoku, KenKen, Kakuro, Futoshiki) with a hierarchical knowledge base to evaluate compositional generalization in AI, enabling fine-grained assessment of reasoning from basic rules to multi-step strategies.

Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence. While few benchmarks exist, many focus on linguistic tasks and lack complex, explicit compositional structures. We introduce ClassicLogic, a new benchmark suite designed to evaluate an agent's ability to learn and compose problem-solving strategies. The benchmark consists of four classic logic puzzles: Sudoku, KenKen, Kakuro, and Futoshiki. Its core innovation is a hierarchical, explicit knowledge base for each game, where complex solving strategies are formally defined as compositions of simpler, foundational strategies. This structure allows for fine-grained evaluation of an agent's reasoning capabilities, from learning basic rules to applying multi-step compositional strategies to solve puzzles of increasing, mathematically validated difficulty. The open-source benchmark provides a challenging new testbed for advancing neuro-symbolic and other advanced AI reasoning systems.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes