A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
Key points
- Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited.
- Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools.
- We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search.
- Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models.
Sources (2)
- [1]A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box OptimizationHugging Face Daily Papers · Oct 8, 12:00 AM
We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol.
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited.
- [2]A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box OptimizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:48 PM · same content
Extractive summary: sentences quoted from the sources.