arXiv cs.CL
7/8/2026

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
Short summary
This study examines whether LLM prompt robustness differs between objective questions (fixed answers) and subjective questions (opinions/values) across four instruction-tuned model families. Using datasets like MMLU, ARC, Political Compass Test, and World Values Survey, the authors apply wording, framing, and format variations to measure answer consistency. Results from binomial GEE analysis show that robustness depends significantly on question type, prompt change category, and model, with a large interaction between dataset type and prompt category.
- •Prompt robustness in LLMs varies significantly between objective and subjective question types
- •Study evaluates four model families across six datasets with multiple prompt variation types
- •Dataset type, prompt category, and model all show significant effects on answer consistency
Generated with AI, which can make mistakes.
Is this a good recommendation for you?