LGCLJun 11

Understanding helpfulness and harmless tension in reward models

arXiv:2606.13209v19.7
Predicted impact top 42% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on multi-objective alignment in language models, this work provides mechanistic insights into why combining helpfulness and harmlessness is challenging.

The paper investigates tension between helpfulness and harmlessness objectives in reward models used for RLHF, finding that mixed-objective models underperform single-objective ones due to interference. Shared neurons disproportionately influence behavior, contributing to alignment tension.

Reward models are a key component of reinforcement learning from human feedback (RLHF), aligning language models toward both helpful and harmless behaviour. However, the internal mechanisms underlying these objectives and their conflicts remain poorly understood. We study alignment tension in reward models trained under helpfulness-only, harmlessness-only, and mixed-objective settings. We find that mixed-objective models often underperform single-objective models, indicating interference between objectives. Using activation-based methods, we identify neurons associated with each objective and study their functional roles via targeted ablations. We find that these neurons causally support their corresponding objectives while often negatively affecting the opposing one. We find that a substantial proportion of neurons are shared between helpfulness and harmlessness, and that these shared neurons exert a disproportionate influence on model behaviour, contributing to alignment tension. Additionally, our results provide insights and mechanistic interpretation into how alignment objectives are represented in reward models and why multi-objective alignment remains challenging, motivating future work on disentangled and controllable alignment methods.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes