AR
arXiv CS.AI
7/9/2026

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics
Short summary
Researchers propose a ReAct-style agentic setup combining LLM reasoning with SageMath feedback and Context7 documentation for solving research-level math problems from the RealMath benchmark. SageMath access improved performance across all evaluated models by an average of 9.7 percentage points, narrowing the gap between open-weight and closed models. GPT-5.5 achieved the highest solve rate of 75.2% with the lowest token usage among tool-enabled configurations.
- •ReAct-style agent combines LLM reasoning with SageMath for computational math
- •SageMath access boosts performance by 9.7pp on average across all models
- •GPT-5.5 achieves 75.2% solve rate with lowest token usage on RealMath benchmark
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
