CLJun 11

LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling

arXiv:2606.12837v115.6
Predicted impact top 64% in CL · last 90 daysOriginality Incremental advance
AI Analysis

Provides a more demanding standard for evaluating long-horizon reasoning and context management in search agents, addressing saturation in existing benchmarks.

LoHoSearch introduces a benchmark of 544 human-verified questions across 11 domains, constructed from a knowledge graph of 7 million Wikipedia entities, where the strongest model achieves only 34.74% accuracy, far below the human difficulty ceiling of prior benchmarks.

Search agent benchmarks exemplified by BrowseComp have rapidly saturated over the past year, with the strongest models surpassing 90% accuracy. Since these benchmarks are predominantly human-authored, annotators lack a global perspective on entity statistics and cannot systematically maximize search space size and structural complexity. This creates a difficulty ceiling that is hard to break. To address this, we introduce LoHoSearch (Long-Horizon Search Agents), a challenging benchmark comprising 544 human-verified questions across 11 domains. LoHoSearch is constructed via an automated pipeline built upon a knowledge graph covering over 7 million Wikipedia entities, which selects relations with large search spaces and assembles them into structurally complex questions with KG-verified unique answers. Our evaluation demonstrates that even the strongest model achieves only 34.74% accuracy, and existing context management strategies (best +6.8%) yield far smaller gains than on prior benchmarks. LoHoSearch provides a more demanding standard for evaluating long-horizon reasoning and context management in search agents.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes