CVJul 28

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

arXiv:2607.254899.2h-index: 1
Predicted impact top 41% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers and clinicians, this review systematically identifies gaps in evaluation and translation of agentic AI, highlighting the need for standardized definitions and prospective validation.

This scoping review maps the emerging landscape of agentic AI in medicine, analyzing 557 studies that use goal-directed agents for clinical tasks. It finds that current evidence is dominated by public benchmarks and simulated settings, with inconsistent evaluation of reliability, safety, and clinical impact.

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes