
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning
Short summary
This paper provides a theoretical framework for in-context search in LLMs, modeling it as approximate inference where the base model defines a prior and self-reflection updates the posterior. The authors prove that when reflections reliably identify early mistakes, in-context search yields exponential improvements over the base model using only polynomial sequential attempts; otherwise, it offers no asymptotic benefit over parallel sampling. They also show these gains are learnable via cross-entropy training on search rollouts and connect the theory to optimal policies under reinforcement learning with verifiable rewards.
- •In-context search modeled as approximate inference: base model is prior, self-reflection is posterior update
- •Exponential gains only when reflections reliably localize early mistakes; otherwise no benefit over parallel sampling
- •Gains are learnable with polynomial sample complexity via cross-entropy training on search rollouts
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
