AICLJul 9

The complexities of patient-centred conversational artificial intelligence

arXiv:2607.0862519.2
Predicted impact top 19% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For developers of health chatbots, this work highlights that ignoring real-world communication diversity can lead to underperformance and health disparities.

The paper analyzes 2,053 real patient-chatbot conversations and develops a patient simulator that models clinical content, emotion, strategy, and communication style. In evaluations, simulated conversations were nearly indistinguishable from real ones (55% human accuracy), and communication style significantly altered triage outcomes across four LLMs.

Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes