Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
For health systems deploying generative AI in patient education, this work reveals that institutional heterogeneity in source documents can undermine the consistency of AI-generated answers, with high-stakes gaps in topics like reproductive health.
This study audits 5.7 million pairwise comparisons across 102 US transplant handbooks to measure institutional heterogeneity in patient education materials, finding that same-center handbooks agree more than same-organ handbooks across centers, information gaps disproportionately affect underrepresented subgroups (e.g., 82% absence for reproductive health), and disagreement is predictable from question framing (AUC=0.77).
Health systems are rapidly deploying generative AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Whether it does depends on a question not previously measured at scale: do the underlying documents themselves agree? We use a structured-output large language model judge to audit 5,730,465 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1,115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) institutional editorial voice statistically transcends organ-type boundaries, with same-center handbooks agreeing across organs more than same-organ handbooks across centers (p = 0.0056); (2) information gaps fall disproportionately on topics central to underrepresented subgroups, with reproductive health showing double jeopardy: it is the single most-silent topic (82% absent) and the highest judge-rated clinical significance when present (86% high-significance disagreements); (3) divergence themes cluster into 991 topics, with immunosuppression and pregnancy timing among the highest-stakes themes; and (4) per-pair disagreement is predictable from question framing alone (AUC = 0.77). We discuss implications for deploying patient-facing generative AI in transplant care.