arXiv cs.CL
7/9/2026

MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning
Short summary
MILES is a framework for self-improving LLM reasoning that dynamically expands step-wise memory with learnable selection heads under realistic test-time constraints. It maintains modular memory units of sub-goal embeddings paired with sub-instructions, using a coarse-to-fine retrieval mechanism that trains selection heads from confident samples and applies them to guide reasoning on uncertain ones. Experiments show it matches or outperforms prior methods with superior accuracy-efficiency tradeoffs.
- •Modular memory with learnable selection heads for test-time LLM reasoning improvement
- •Coarse-to-fine retrieval: expand memory from confident samples, guide uncertain ones
- •Matches or outperforms prior methods with better accuracy-efficiency tradeoffs
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
