Back to feed
arXiv cs.CL
arXiv cs.CL
7/8/2026
Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Short summary

This study examines whether LLM prompt robustness differs between objective questions (fixed answers) and subjective questions (opinions/values) across four instruction-tuned model families. Using datasets like MMLU, ARC, Political Compass Test, and World Values Survey, the authors apply wording, framing, and format variations to measure answer consistency. Results from binomial GEE analysis show that robustness depends significantly on question type, prompt change category, and model, with a large interaction between dataset type and prompt category.

  • Prompt robustness in LLMs varies significantly between objective and subjective question types
  • Study evaluates four model families across six datasets with multiple prompt variation types
  • Dataset type, prompt category, and model all show significant effects on answer consistency

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more