CLCYLGJun 27

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

arXiv:2606.2896322.9
Predicted impact top 12% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers using LLMs to simulate social surveys, this work identifies systematic biases and evaluates methods to improve fidelity, though findings are incremental.

The paper investigates whether LLMs can recover population-level statistical characteristics from small pilot samples, proposing a three-axis fidelity framework (structural, marginal, individual). Fine-tuning on small pilot data achieves balanced fidelity but varies across subsamples, threatening pluralistic alignment.

Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity. Using a COVID-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine-tuning. The findings suggest that fine-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes