AIAug 3

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

arXiv:2608.0216322.9
Predicted impact top 7% in AI · last 90 daysOriginality Synthesis-oriented
AI Analysis

Provides a new benchmark for evaluating deep research capabilities in AI systems, addressing the need for verifiable and automatically constructed evaluation tasks.

The paper introduces a verifiable benchmark of 500 deep research tasks across 31 topics, constructed automatically via an iterative pipeline. Experiments show the benchmark discriminates among models and query types, and its rubrics enable fine-grained, human-aligned evaluation.

Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existing human-authored materials, while fully automatic construction struggles to ensure consistent and traceable verification. To address this gap, we introduce a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research. The benchmark is constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline that progressively transforms simple questions into deep research tasks. Each task is represented as a directed acyclic graph (DAG) of atomic steps and associated checkpoints, enabling the query, DAG, and rubrics to evolve together in a controlled manner. Experiments demonstrate that the benchmark clearly discriminates among models and query types, while its fact-grounded pointwise rubrics enable fine-grained, human-aligned, and stable evaluation. Our data, implementation, and results are publicly available.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes