AIJul 2

Scaling Trends for Lie Detector Oversight in Preference Learning

arXiv:2607.0156717.4
Predicted impact top 25% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For AI safety researchers, this work shows that lie detector oversight scales favorably with model size, but highlights a critical sensitivity to distribution shift that limits practical deployment.

Scaling SOLiD to larger models reduces undetected deception from 34% (1B) to 14% (405B) at 99% true positive rate, and human labelers can be removed from fine-tuning without significant increase in deception. However, distribution shift can cause impractical false positive rates.

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes