CLIRJun 15

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

arXiv:2606.1704115.81 citations
Predicted impact top 62% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers evaluating LLM-based scientific reasoning, this work provides a benchmark with ground truth across the full meta-analysis pipeline and identifies the screening step as the primary failure point.

MetaSyn, a dataset of 442 expert-curated meta-analyses from Nature Portfolio, reveals that current LLM agents achieve at most 52.7% recall of ground-truth included literature despite a retrieval ceiling of 90.9% recall at K=200, highlighting a critical screening bottleneck.

Meta-analysis is a demanding form of evidence synthesis that combines literature retrieval, PI/ECO-guided study selection, and statistical aggregation. Its structured, verifiable workflow makes it an ideal substrate for evaluating systematic scientific reasoning, yet existing benchmarks lack ground truth across the full retrieval-screening-synthesis pipeline. We introduce MetaSyn, a dataset of 442 expert-curated meta-analyses from Nature Portfolio journals. Each entry pairs a research question with PI/ECO criteria, a retrieval corpus of 140k PubMed articles, verified positive studies, hard negatives that are topically similar but PI/ECO-ineligible, and complete search strategies and date bounds. Benchmarking twelve pipeline configurations (nine RAG variants and a protocol-driven agent) reveals a critical screening bottleneck: despite a retrieval ceiling of 90.9% recall at K=200, no system recovers more than 52.7% of ground-truth included literature. Current LLMs fail to reliably separate eligible studies from PI/ECO-failing distractors in pools of comparable topical relevance. Stage-attributed metrics capture where systems succeed and fail; a single end-to-end score does not.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes