Back to feed
arXiv cs.CL
arXiv cs.CL
8/5/2026
BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

Short summary

BBOWP-Bench introduces a novel benchmark for evaluating LLMs on black-box optimization word problems, where systems must infer both a search space and an optimization algorithm from natural-language descriptions. The benchmark includes executable evaluation environments and human-designed baselines. Results show current LLMs can select suitable algorithms based on evaluation budgets but struggle with search-space design when problem descriptions are less informative.

  • New benchmark for LLM-based black-box optimization problem formulation from natural language
  • Evaluates both search-space design and algorithm selection with executable environments
  • LLMs handle algorithm selection well but struggle with identifying key variables and balancing ranges

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more