SEAIJun 16

Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations

arXiv:2606.175883.4
Predicted impact top 90% in SE · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers conducting systematic reviews, this work provides a qualitative understanding of LLM failures and practical recommendations, but it is incremental as it does not propose a novel method or achieve SOTA results.

The study qualitatively investigates why large language models (LLMs) disagree with human experts in title-abstract screening for systematic reviews, identifying causes like boundary ambiguity and keyword overemphasization. It proposes actionable recommendations such as validating semantic understanding and using multiple LLMs, though future validation is needed.

Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting mixed accuracy. However, questions of reliability remain largely unaddressed. In this study, we go beyond quantitative LLM-human agreement metrics and qualitatively investigate how and why LLMs fail. We also propose actionable recommendations. We analyzed disagreements between LLMs and researchers across six software engineering SRs and over 1,000 primary study papers. For each SR, papers were screened independently by human experts and LLMs in zero-shot mode, resulting in Kappa values ranging from 0.52 to 0.77. Qualitative analysis suggests that human-LLM disagreement results from recurring, identifiable causes, such as boundary ambiguity in key terms, keyword overemphasization, and incorrect topic inference. Based on these findings, we propose recommendations such as validating semantic understanding before deployment, running multiple LLMs, and focusing validation efforts on borderline cases. Future studies are needed to validate the impact of our recommendations, and community efforts are needed to develop normative guidelines on LLM usage in SRs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes