Beyond Accuracy: Interpreting Topic Representation in Suicide Ideation Detection Models
For mental health AI safety, this provides insight into model interpretability beyond accuracy, though the analysis is exploratory and domain-specific.
This work analyzes how suicide ideation detection models internally represent psychological risk factors, showing that topic-aware augmentation improves the clarity and distinctness of underrepresented factors such as immigration, family issues, and financial crisis.
Suicide ideation detection models are typically evaluated using aggregate performance metrics, yet little is known about how they internally represent psychologically meaningful risk factors. In high-stakes mental health applications, understanding these internal representations is essential for safety, transparency, and responsible deployment. In this work, we move beyond accuracy and analyze how suicide detection models trained on original and topic-augmented datasets encode psychological risk factors in their internal representation space. Using visualization and geometric analysis, we examine the coherence and separability of topic-related features. Our results show that topic-aware augmentation increases the clarity and distinctness of underrepresented psychosocial risk factors such as immigration, family issues, and financial crisis. These findings suggest that augmentation not only improves model performance but also leads to more structured and interpretable internal representations.