On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.
Key points
- We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance.
- We investigate two complementary classes of interventions: harness-based mechanisms that expose timing information and enforce deadlines, and reinforcement learning with budget-aware rewards.
- RL-trained policies learn when to stop but often fill extra time with repeated actions, and GRPO training on multiple budgets tends to collapse toward the strategy learned for the shortest budget.
- Our results reveal a gap between time adherence and productive time allocation, which remains a central challenge for budget-conditioned agents.
Sources (1)
- [1]On the Clock: Towards Punctual and Productive Time-Budgeted AI AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:35 PM
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively.
We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance.
Extractive summary: sentences quoted from the sources.