AIJul 1

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

arXiv:2608.21374h-index: 1Has Code
Originality Highly original
AI Analysis

This work addresses the critical problem of rigorously evaluating automatically generated literature reviews, which is essential for researchers and the scientific community, by providing a human-centric evaluation platform and a more aligned automated judge.

This paper introduces LitReview Arena, a battle-style evaluation platform for literature review agents, collecting 3k expert judgments. It found that current systems win only 23.0% of decisive matches against human drafts, while agentic LLMs outperform base language models by over 60%. The paper also developed LitJudge, an expert-calibrated evaluator, which improved alignment with human experts to Spearman's rho=0.78.

Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes