AIJun 22

HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs

arXiv:2606.2323814.5Has Code
Predicted impact top 43% in AI · last 90 daysOriginality Highly original
AI Analysis

For AI safety and legal/financial applications, this benchmark exposes that current LLMs cannot handle higher-order reasoning, which is critical for verifiable decision-making.

LLMs struggle with higher-order logical reasoning, achieving only 50.64% average accuracy on the new HOLMES benchmark, with the best model reaching 59.54%, revealing a key bottleneck for reliable AI.

Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level deduction over fixed predicates. This misses many realistic scenarios where models must reason over rules, predicates, functions, constraints, and decision procedures themselves. We introduce HOLMES (Higher-Order Logic Meets real-world Explainable Symbolic reasoning), the first real-world benchmark for higher-order symbolic reasoning in LLMs, containing 1379 instances. Built on higher-order logic, HOLMES pairs natural-language problems with HOL formalizations, ground-truth answers, verifiable reasoning traces, and fine-grained controllable reasoning factors across law and finance. Experiments show that current LLMs still struggle on HOLMES, with an average accuracy of only 50.64% and the best model reaching 59.54%. Our analyses further reveal that high final-answer accuracy can mask shortcut reasoning in conflict-resolution settings, while performance drops sharply under scope-conditioned and compositional reasoning. These findings identify higher-order symbolic reasoning as a key bottleneck for building reliable and verifiable LLMs. The project code and dataset are publicly available at https://github.com/wuyucheng2002/HOLMES.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes