Back to feed
arXiv cs.CL
arXiv cs.CL
7/21/2026
Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM

Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM

Short summary

This study shows that LLMs sometimes commit to an answer before reasoning and then justify it, even when the answer contradicts task premises. Using a car-wash probe on Qwen3-8B, wrong commitments occurred in 85-100% of rollouts across conditions. Preliminary activation-level probing reveals hidden states lean toward the pre-committed answer before it is emitted, suggesting reasoning may be post-hoc rationalization.

  • Qwen3-8B commits to wrong answers in 85-100% of rollouts on a simple reasoning probe, even with extended thinking budgets
  • Activation-level probing shows hidden states encode the pre-committed answer before it is emitted
  • Findings suggest chain-of-thought reasoning can function as post-hoc justification rather than genuine derivation

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more