arXiv cs.CL
7/21/2026

Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM
Short summary
This study shows that LLMs sometimes commit to an answer before reasoning and then justify it, even when the answer contradicts task premises. Using a car-wash probe on Qwen3-8B, wrong commitments occurred in 85-100% of rollouts across conditions. Preliminary activation-level probing reveals hidden states lean toward the pre-committed answer before it is emitted, suggesting reasoning may be post-hoc rationalization.
- •Qwen3-8B commits to wrong answers in 85-100% of rollouts on a simple reasoning probe, even with extended thinking budgets
- •Activation-level probing shows hidden states encode the pre-committed answer before it is emitted
- •Findings suggest chain-of-thought reasoning can function as post-hoc justification rather than genuine derivation
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
