AR
arXiv CS.AI
7/16/2026

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
Short summary
This paper introduces interventional grounding audits, a black-box method that tests whether LLM chain-of-thought reasoning genuinely depends on its stated premises by substituting predicates with fresh symbols and re-running the model. On ProntoQA with GPT-4o, the method achieves F1=0.806 for detecting proof-tree dependencies, far outperforming self-consistency baselines (F1=0.343). The authors find 66% of correctly-solved problems contain at least one reasoning step insensitive to a direct proof-tree dependency—a 'right answer, wrong reasoning' signal.
- •Black-box predicate substitution audits LLM CoT premise dependency
- •Achieves F1=0.806 vs 0.343 for self-consistency baseline on GPT-4o
- •66% of correct answers contain at least one insensitive reasoning step
Generated with AI, which can make mistakes.
Is this a good recommendation for you?