CYJun 13

Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks

arXiv:2607.22606
Originality Incremental advance
AI Analysis

For health systems deploying generative AI in patient education, this work reveals that institutional heterogeneity in source documents can undermine the consistency of AI-generated answers, with high-stakes gaps in topics like reproductive health.

This study audits 5.7 million pairwise comparisons across 102 US transplant handbooks to measure institutional heterogeneity in patient education materials, finding that same-center handbooks agree more than same-organ handbooks across centers, information gaps disproportionately affect underrepresented subgroups (e.g., 82% absence for reproductive health), and disagreement is predictable from question framing (AUC=0.77).

Health systems are rapidly deploying generative AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Whether it does depends on a question not previously measured at scale: do the underlying documents themselves agree? We use a structured-output large language model judge to audit 5,730,465 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1,115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) institutional editorial voice statistically transcends organ-type boundaries, with same-center handbooks agreeing across organs more than same-organ handbooks across centers (p = 0.0056); (2) information gaps fall disproportionately on topics central to underrepresented subgroups, with reproductive health showing double jeopardy: it is the single most-silent topic (82% absent) and the highest judge-rated clinical significance when present (86% high-significance disagreements); (3) divergence themes cluster into 991 topics, with immunosuppression and pregnancy timing among the highest-stakes themes; and (4) per-pair disagreement is predictable from question framing alone (AUC = 0.77). We discuss implications for deploying patient-facing generative AI in transplant care.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes