CLJun 24

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

arXiv:2606.2544224.4Has Code
Predicted impact top 21% in CL · last 90 daysOriginality Incremental advance
AI Analysis

It addresses the practical problem of aligning LLMs with rapidly evolving safety policies in real-world deployment, where conventional data-driven methods are costly or delayed.

PolicyAlign directly aligns LLMs with natural-language safety policies without requiring costly supervision data, achieving consistent safety improvements across multiple models while maintaining low over-refusal and preserving general capabilities.

Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. However, in real-world deployment, emerging safety requirements are often specified as natural-language policies, while corresponding supervision data may be costly, delayed, or unavailable. This creates a mismatch between rapidly evolving safety policies and conventional data-driven alignment methods. To address this, we propose PolicyAlign, a simple yet effective framework for directly aligning LLMs with safety policies. Given a safety policy, PolicyAlign first synthesizes policy-violating instructions and then performs on-policy self-distillation to internalize policy-guided behavior. To improve training stability and data efficiency, we further introduce Policy-Sensitive Filtering, which selects instructions where the policy induces the largest behavioral shift. Experiments across multiple models show that PolicyAlign consistently improves safety while maintaining low over-refusal and preserving general capabilities. PolicyAlign also generalizes to medical, legal, and financial safety scenarios, highlighting its potential as a scalable and maintainable approach to policy-based LLM safety alignment. The code is released at https://github.com/Qwen-Applications/PolicyAlign.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes