arXiv cs.CL
7/28/2026

Not All LLM Reasoning is Visible in the Chain-of-Thought
Short summary
Researchers demonstrate that frontier LLMs perform invisible reasoning using semantically irrelevant filler tokens, achieving up to 13 percentage point accuracy improvements across 13 models. Claude Opus 4.5 can satisfy hidden modular arithmetic constraints via filler tokens without sacrificing primary task accuracy. RL gives Qwen3-235B preferences over filler content, but neither RL nor SFT produces persistent test-time benefits, indicating frontier models already compute without interpretable traces.
- •Frontier LLMs use filler tokens for invisible reasoning, gaining up to 13pp accuracy improvements
- •Claude Opus 4.5 satisfies hidden arithmetic constraints via filler tokens invisible to CoT monitoring
- •RL and SFT do not produce persistent test-time filler token benefits, but models already compute without interpretable traces
Generated with AI, which can make mistakes.
Is this a good recommendation for you?