Yubo Li, Rema Padman, Ramayya Krishnan
Health systems are rapidly deploying generative AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Whether it does depends on a question not previously measured at scale: do the underlying documents themselves agree? We use a structured-output large language model judge to audit 5,730,465 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1,115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) institutional editorial voice statistically transcends organ-type boundaries, with same-center handbooks agreeing across organs more than same-organ handbooks across centers (p = 0.0056); (2) information gaps fall disproportionately on topics central to underrepresented subgroups, with reproductive health showing double jeopardy: it is the single most-silent topic (82% absent) and the highest judge-rated clinical significance when present (86% high-significance disagreements); (3) divergence themes cluster into 991 topics, with immunosuppression and pregnancy timing among the highest-stakes themes; and (4) per-pair disagreement is predictable from question framing alone (AUC = 0.77). We discuss implications for deploying patient-facing generative AI in transplant care.