AIJul 23

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

arXiv:2607.2092619.9
Predicted impact top 15% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers developing scientific AI agents, this benchmark exposes critical gaps in current models' ability to handle realistic multi-step scientific workflows, highlighting the need for improved reasoning and information integration.

SciExplore is a benchmark for evaluating LLMs and agents on scientific information-seeking and reasoning tasks across four complexity levels. Testing over ten state-of-the-art models reveals sharp performance degradation with task complexity, with near-zero accuracy on the most challenging structured synthesis tasks.

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes