AIAPAug 1

Large language models improve physician accuracy but lead to false reliance

arXiv:2608.0081710.6
Predicted impact top 57% in AI · last 90 daysOriginality Incremental advance
AI Analysis

This paper addresses the critical issue of over-reliance on AI in clinical decision-making, showing that while source-linked LLMs improve accuracy, they also introduce a dangerous tendency to accept incorrect advice when citations are present.

The authors developed CORA, a retrieval-augmented LLM, and tested it with 46 physicians, finding that accuracy improved from 70.8% to 82.6% with assistance, but citation-supported incorrect answers reduced physician resistance from 92% to 34.8%, revealing a safety risk.

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes