arXiv cs.CL
8/5/2026

BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems
Short summary
BBOWP-Bench introduces a novel benchmark for evaluating LLMs on black-box optimization word problems, where systems must infer both a search space and an optimization algorithm from natural-language descriptions. The benchmark includes executable evaluation environments and human-designed baselines. Results show current LLMs can select suitable algorithms based on evaluation budgets but struggle with search-space design when problem descriptions are less informative.
- •New benchmark for LLM-based black-box optimization problem formulation from natural language
- •Evaluates both search-space design and algorithm selection with executable environments
- •LLMs handle algorithm selection well but struggle with identifying key variables and balancing ranges
Generated with AI, which can make mistakes.
Is this a good recommendation for you?