CLAIJul 6

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

arXiv:2607.0555412.1
Predicted impact top 70% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers using LLMs for survey-style evaluations, this work highlights that prompt robustness is task-dependent, cautioning against treating responses as stable measures of values or beliefs.

The study investigates whether prompt robustness in LLMs differs between objective and subjective questions, finding that robustness depends on question type, prompt change, and model, with significant interactions between dataset type and prompt category.

Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes