Back to feed
arXiv cs.CL
arXiv cs.CL
7/8/2026
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

Short summary

This study introduces a psychometric battery that separates logical verdicts from surface-level framing artifacts in LLM yes/no moral judgments. Frontier models show nearly format-invariant moral stances, but forcing binary verdicts overlays decomposable biases: an order bias toward the last-printed option and a lexical pull toward the word "no," substantial only in Claude models and near-zero for GPT-5.5 and Gemini. The findings show models are not inherently drawn toward rejecting—apparent moral shifts follow printed surface form, not the verdict it carries.

  • LLM yes/no moral judgment shifts are framing artifacts, not genuine moral stance changes
  • Claude models show substantial order and lexical biases (-0.32 to -0.86); GPT-5.5 and Gemini show near-zero bias
  • A minimal model P = σ((θ ± m)/s) captures framing susceptibility and moral decisiveness, applicable to any dilemma set

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more