Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings
For NLP researchers working on social meaning and bias, this work provides a method to capture annotator diversity, though the gains are incremental.
The paper addresses the perspectival nature of social meaning in language by modeling how interpretations vary across demographic groups. Their proposed fusion embeddings, integrating textual and demographic representations, achieve consistent improvements of +5.9-6.5% relative macro PR-AUC over text-only baselines.
Social meaning in language is inherently perspectival, varying across annotator backgrounds, demographics, and ideological positions. However, most NLP systems collapse this variation into a single ground-truth label, ignoring the diversity of interpretations. In this work, we model social dimensions along a perspectivist spectrum, capturing how interpretations vary across demographic groups on a dataset consisting of 28k human annotations. We benchmark multiple modeling paradigms, including zero-shot, few-shot, and fine-tuned approaches, and propose fusion embeddings that integrate textual and demographic representations. Our fusion models yield consistent and statistically significant improvements over text-only baselines across all fusion strategies (+5.9-6.5% relative macro PR-AUC), with shuffle ablations confirming that demographic profiles carry genuine predictive signal rather than spurious correlations.