AIJun 19

AI Alignment From Social Choice Perspectives

arXiv:2606.2155017.4
Predicted impact top 30% in AI · last 90 daysOriginality Synthesis-oriented
AI Analysis

For AI alignment researchers, it provides a structured lens to address value conflicts in human feedback aggregation.

The paper surveys how social choice theory can inform the aggregation of conflicting human feedback in AI alignment, identifying failure modes and expanding design options for handling disagreement.

Alignment from human feedback uses human judgments about model outputs to steer the behavior of language models after pretraining. When those judgments reflect conflicting views of desirable behavior, the learned objective becomes an aggregate determination of what the model should prefer. We survey recent work that has studied this aggregation problem through the lens of social choice theory. We illustrate how the social choice perspective helps identify failure modes in the feedback aggregation layer and reveals a broader design space for handling disagreement in explicit and principled ways.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes