CLJun 14

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

arXiv:2606.1551717.0
Predicted impact top 55% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For LLM developers, SHARD offers a method to reduce refusal and improve helpfulness on sensitive queries without sacrificing safety, though gains are incremental over existing alignment techniques.

SHARD improves safe-helpfulness of LLMs by self-reframing sensitive prompts and responses, achieving competitive performance with teacher distillation while preserving safety across DNA and LINGUASAFE benchmarks.

Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational needs that can be answered safely. We introduce SHARD, a self-reframing distillation method to improve safe-helpfulness. It first rewrites sensitive prompts to surface benign intent using philosophical guidelines, then reframes its original responses into safe, more helpful ones, and finally fine-tunes the model on its self-reframed responses. Across DNA and the English subset of LINGUASAFE, SHARD improves helpfulness for most model families while preserving safety. It also remains competitive with distillation from a larger teacher model, suggesting that models can internalize safe and helpful behavior elicited from their own. Warning: This paper contains content that may be offensive or harmful.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes