Affective AI Safety: The Missing Piece in LLM Safety
For AI safety researchers and regulators, the paper highlights an undertheorized class of harms, but it is primarily a conceptual framework without empirical validation.
The paper identifies a gap in AI safety research regarding emotional harms from LLMs, proposes a taxonomy of affective harms (self-alienation, bias, relational), and argues that current frameworks inadequately address these issues, requiring dedicated technical and regulatory approaches.
AI safety research has focused predominantly on epistemic and physical harms (e.g., misinformation, bias, system reliability) while the risks that arise from AI systems' engagement with human emotional life have remained fragmented and undertheorised. We propose affective safety as a unified class of AI safety concerns grounded in the fact that humans are affective beings. We develop a taxonomy of affective harms and identify recurring harm types: (1) affective self-alienation, (2) fairness and bias harms, and (3) relational harms. We show that their recurrence across system types reflects structural properties of how AI systems engage with human emotion and survey the current safety landscape and show that existing frameworks address affective safety either narrowly or not at all. We conclude by identifying the technical and regulatory challenges specific to this class of harms and argue that affective safety requires dedicated frameworks that engage with cumulative, relational, and identity-level effects.