AIJul 16

Transcoders for Investigating Deception in Language Models

arXiv:2607.1479111.4h-index: 2
Predicted impact top 13% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For AI safety researchers, it demonstrates a method for detecting and understanding deceptive behavior in language models, though the approach is incremental.

The paper uses transcoders to analyze deceptive behavior in language models, identifying deception-related features that exert stronger influence on deceptive outputs and enable circuit-level analysis.

Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour. In this paper, we investigate the use of transcoders to analyse deceptive behaviour in language models, a behaviour that poses a safety and security risk. Using a Qwen3-4B model with pre-trained transcoders, specifically per-layer transcoders (PLTs), we construct attribution graphs that capture feature activations and inter-feature dependencies, allowing circuit-level analysis of deception. Through feature steering and circuit analysis, we identified a dictionary of deception-related features and show that these features exert a stronger influence on deceptive outputs, as they produce predictable shifts between deceptive and non-deceptive responses. These findings suggest that deception emerges from internal model mechanisms and highlight the potential of transcoders for behavioural monitoring and early detection of security vulnerabilities related to malicious behaviours in language models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes