AIJun 16

Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games

arXiv:2606.1765720.0
Predicted impact top 21% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers using LLMs for human behavior simulation in safety evaluations and training, this work provides a method to incorporate cognitive models, though the gains are incremental and domain-specific.

The authors propose Equation-to-Behavior Prompting and Equation-to-Behavior RL to make large language models simulate cognitive models of human decision-making in persuasion games, finding that large models can approximate equation-based specifications via prompting while small models improve with RL, reducing belief error by 26.5% in out-of-distribution settings and improving training diversity by 2.5%–12% over Bayesian-only training.

People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes