arXiv cs.CL
7/8/2026

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
Short summary
This study introduces a psychometric battery that separates logical verdicts from surface-level framing artifacts in LLM yes/no moral judgments. Frontier models show nearly format-invariant moral stances, but forcing binary verdicts overlays decomposable biases: an order bias toward the last-printed option and a lexical pull toward the word "no," substantial only in Claude models and near-zero for GPT-5.5 and Gemini. The findings show models are not inherently drawn toward rejecting—apparent moral shifts follow printed surface form, not the verdict it carries.
- •LLM yes/no moral judgment shifts are framing artifacts, not genuine moral stance changes
- •Claude models show substantial order and lexical biases (-0.32 to -0.86); GPT-5.5 and Gemini show near-zero bias
- •A minimal model P = σ((θ ± m)/s) captures framing susceptibility and moral decisiveness, applicable to any dilemma set
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
