AIJun 27

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

arXiv:2606.287709.5
Predicted impact top 60% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners seeking fine-grained, interpretable control over LLM personality without prompt engineering or fine-tuning, this work offers a novel mechanistic approach.

The paper proposes a mechanistic interpretability method that intervenes on LLM latent features to steer OCEAN personality traits, using sparse autoencoders and contrastive activation analysis to identify and shift latent directions, achieving controllable personality expression while preserving language modeling performance.

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes