ResearchPod Summary
Researchers often use survey-style evaluations to infer the political values, social attitudes, or beliefs of large language models (LLMs). However, these evaluations frequently rely on a single prompt, assuming that the model's response is a stable reflection of its internal state. This paper investigates whether this assumption holds by comparing prompt robustness across two distinct types of tasks: Type-I (objective, multiple-choice questions with fixed answers) and Type-II (subjective, opinion-based survey items).
The authors evaluated four instruction-tuned model families (Gemma, Llama, Mistral, and Qwen) across three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). They applied a taxonomy of seven perturbation categories—including paraphrasing, spelling noise, label substitution, and option shuffling—to measure how consistently models answer the same question when the prompt is varied without changing the underlying task intent.
The study reveals that prompt robustness is not a universal model property but is instead highly dependent on the interaction between the model, the dataset, and the specific type of prompt change. Subjective datasets consistently demonstrate lower response consistency than objective ones, with an average instability gap of approximately 6 percentage points.
Among the perturbation categories, option-order shuffling is the most significant driver of instability, causing a dramatic drop in consistency across both task types. While models are relatively robust to semantic changes like paraphrasing or lexical substitution, they are highly sensitive to how answer choices are presented. The statistical analysis confirms that the interaction between dataset type and prompt category is significant, indicating that subjective questions are uniquely vulnerable to surface-level formatting changes.
These findings challenge the validity of using single-prompt survey evaluations to measure LLM "beliefs." Because subjective responses are so easily shifted by minor changes in formatting or option order, a single answer cannot be reliably interpreted as a stable model trait. The authors argue that researchers must move away from single-prompt evaluations and instead report consistency scores across a range of perturbation categories to distinguish between genuine model stances and artifacts of prompt design.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.