HCJun 15

A comparison of human and LLM-simulated participants in a writing style task

arXiv:2606.1677812.3
Predicted impact top 15% in HC · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers considering LLM simulations as human replacements, this paper provides a cautionary case study highlighting methodological challenges and discrepancies.

This study compares LLM-simulated participants (GPT-4o) with 30 human participants in a writing style preference inference task, finding that the LLM exhibits bias and lacks depth, failing to accurately simulate human behavior.

Because large language models (LLMs) can produce natural language that is sometimes indistinguishable from texts produced by people, some researchers are starting to consider replacing human participants with LLM simulations. In this study, we test the extent to which the findings of a simulation with an LLM prompted to act as a synthetic participant match those obtained from 30 human participants. In our experiments, we evaluated how well writing style preference inference algorithms adapted to a participant over repeated interactions, compared to a baseline. We discover hints of bias and a lack of depth in GPT-4o's text generation and judgement that prevent it from accurately simulating people's behavior. Our results also hint at human biases that highlight the importance of considering human factors in the evaluation of systems that depend on human-automation interaction. Rather than treating these discrepancies as evidence for or against the validity of LLM-simulated participants, we present this study as a case analysis of methodological and design challenges.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes