AIJun 5

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

arXiv:2606.0787417.7h-index: 3
Predicted impact top 29% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners relying on LLMs for safety evaluation, this work highlights a critical limitation in the flexibility of current evaluators.

The paper investigates whether LLM-as-judge evaluators can adapt their safety assessments based on new in-context information or differing safety definitions, finding that they are largely unable to adjust when context contradicts their internal priors.

LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs-as-judges: their susceptibility to relying on in context-information, and their steerability to differing safety definitions, which may not align with their internal safety priors. We evaluate the safety judging abilities of many generalist LLMs and safety-specific judges, and investigate the impact of task demonstrations, novel in-context information, and changing safety definitions. We find that while LLM-judges can learn from new information, they are broadly unlikely to adjust their evaluations if the context or safety definition contradicts their prior.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes