CLJun 30

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

arXiv:2606.3160818.2
Predicted impact top 33% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners deploying LLMs in clinical settings, CLExEval exposes critical evaluation illusions and failure patterns that standalone automated metrics miss, highlighting the need for expert-grounded validation.

CLExEval introduces a human-in-the-loop framework to evaluate LLM clinical reasoning under progressive information masking, revealing that GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity and that 68.6% of reasoning traces contain correct diagnoses not reflected in final answers, while LLM-as-a-Judge evaluations overestimate reliability without expert validation.

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce CLExEval, a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking. CLExEval combines 5,600 expert-physician annotations with 200 clinical reasoning traces derived from 40 rare diagnostic cases. Our analysis identifies three recurring failure patterns: (i) verbosity bias, where GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity; (ii) a hidden knowledge paradox, where a specialist model reaches 92.5% maximum diagnostic potential but fails to retrieve that knowledge reliably in verbose contexts; and (iii) a 68.6% reasoning-to-output mismatch, where correct diagnoses appear in reasoning traces but are not reflected in final answers. We further evaluate the LLM-as-a-Judge paradigm on a human-verified failure set (n = 142). GPT-4o-mini approved 47.9% of clinically incorrect outputs, while HuatuoGPT-o1 approved all validly scored failures and showed a positive self-preference bias. These results suggest that standalone automated clinical evaluations can substantially overestimate clinical reliability without expert-grounded validation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes