Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text
For clinical NLP researchers, the paper highlights how dataset construction assumptions can distort the interpretation of suicidality detection models, urging caution in treating labels as ground truth.
The paper argues that EHR-based suicidality datasets encode specific operationalizations of suicidality shaped by data construction choices, using the ScAN dataset as a case study. It shows that identical labels subsume heterogeneous clinical framings differing in temporality, negation, and uncertainty.
Clinical NLP increasingly relies on electronic health record (EHR) data to detect suicidal behaviors, treating clinical documentation as more reliable ground truth than social media. We argue that this framing obscures how EHR-based suicidality datasets encode a particular operationalization of suicidality, shaped by who authors the data, how episodes are bounded, and how ambiguity is resolved. We ground this argument in a case study of the ScAN dataset, built over MIMIC-III clinical notes. We show how governance constraints, ICD-based cohort selection, single-annotator labeling, and hospital-stay-level aggregation produce labels that reflect clinician-documented judgments, treat suicidality as a bounded episode, and assume that intent can be reliably inferred from documentation. A linguistic analysis demonstrates that identical labels subsume heterogeneous clinical framings differing in temporality, negation, and uncertainty. We argue that clinical NLP should examine the assumptions embedded in suicidality datasets before interpreting their labels as ground truth.