Back to feed
arXiv cs.CL
arXiv cs.CL
7/15/2026
MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization

MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization

Short summary

MAGE is a controlled analysis framework for studying how components of iterative prompt optimization interact, revealing the Prompt Optimization Coupling Effect (POCE): multiple stochastic optimization signals in a closed loop simultaneously improve performance and amplify variance. Key findings include that failure-grounded reflection is essential (score-only or abstract-critique methods fail), MAGE outperforms GEMA 46.4% vs 34.0% on GSM8K-Hard, and in low-data regimes well-designed fixed prompts beat all reflective optimizers. The work suggests prompt optimization systems should be evaluated on both performance and stability, not just peak accuracy.

  • Discovers POCE: combining prompt optimization signals improves accuracy but amplifies variance
  • MAGE achieves 46.4% vs GEPA's 34.0% on GSM8K-Hard with comparable variance
  • In low-data regimes (N=30), fixed prompts outperform all reflective optimizers

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more