
MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization
Short summary
MAGE is a controlled analysis framework for studying how components of iterative prompt optimization interact, revealing the Prompt Optimization Coupling Effect (POCE): multiple stochastic optimization signals in a closed loop simultaneously improve performance and amplify variance. Key findings include that failure-grounded reflection is essential (score-only or abstract-critique methods fail), MAGE outperforms GEMA 46.4% vs 34.0% on GSM8K-Hard, and in low-data regimes well-designed fixed prompts beat all reflective optimizers. The work suggests prompt optimization systems should be evaluated on both performance and stability, not just peak accuracy.
- •Discovers POCE: combining prompt optimization signals improves accuracy but amplifies variance
- •MAGE achieves 46.4% vs GEPA's 34.0% on GSM8K-Hard with comparable variance
- •In low-data regimes (N=30), fixed prompts outperform all reflective optimizers
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
