CLJul 1

Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

arXiv:2607.0110320.1
Predicted impact top 25% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For medical AI benchmarking, this work highlights that automated LLM evaluators can match physician agreement but fail to replicate clinical metacognition and independence, posing risks for deployment.

The paper introduces MedQADE, the first standardized open-response clinical benchmark for German, and finds that while the top LLM evaluator (Gemini 3 Flash) achieves agreement with physicians (kappa=0.694 vs. 0.709), it lacks clinical caution by never abstaining and exhibits lineage-dependent biases. This shows statistical alignment does not ensure clinical safety.

Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large Language Model (LLM) evaluators. The top-performing evaluator model, Gemini 3 Flash, reached alignment consistent with the physician ceiling (\k{appa} = 0.694 vs. \k{appa} = 0.709), though wide confidence intervals limit interpretation. Despite this statistical alignment, automated evaluators exhibited near-absent clinical metacognition: physicians scaled abstention with item difficulty, while frontier models assigned definitive scores in every case. We additionally quantified systematic lineage-dependent biases, where models preferentially scored architectural siblings, an effect independent of language. These results show that statistical alignment does not ensure clinical caution, and that evaluator independence requires explicit verification.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes