AIAug 7

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

CMU
arXiv:2608.0779624.7h-index: 8
Predicted impact top 4% in AI · last 90 daysOriginality Highly original
AI Analysis

This benchmark addresses the critical need for reliable and defensible clinical reasoning in large language models for healthcare professionals, by evaluating their ability to investigate real longitudinal EHR data and ground conclusions in verifiable evidence.

This paper introduces CliniCARE-Bench, a benchmark for evaluating large language models in retrospective clinical audit over electronic health records. It features 25 clinician-validated scenarios across 750 patient-specific cases from real MIMIC-IV data. The benchmark reveals that while raw accuracy for 16 agentic systems ranges from 65.3-76.1%, defect-free accuracy, which penalizes shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes