Back to feed
AR
arXiv CS.AI
7/9/2026
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

Short summary

Researchers propose a ReAct-style agentic setup combining LLM reasoning with SageMath feedback and Context7 documentation for solving research-level math problems from the RealMath benchmark. SageMath access improved performance across all evaluated models by an average of 9.7 percentage points, narrowing the gap between open-weight and closed models. GPT-5.5 achieved the highest solve rate of 75.2% with the lowest token usage among tool-enabled configurations.

  • ReAct-style agent combines LLM reasoning with SageMath for computational math
  • SageMath access boosts performance by 9.7pp on average across all models
  • GPT-5.5 achieves 75.2% solve rate with lowest token usage on RealMath benchmark

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more