Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets.
ProofPaper ↗
Key points
- Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive.
- Can LLM agents autonomously create cheaper solutions for such workloads?
- We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost.
- Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities.
Sources (1)
- [1]Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:57 PM
We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets.
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Sep 28, 2026Notes on NVIDIA Nemotron
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0