AION
Research paperLarge Language Models · Reinforcement Learning · Robotics & Embodied AI1 source · Oct 6, 2026

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks.

Key points

  • Large language model agents are increasingly deployed to perform complex tasks in real-world environments.
  • Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience.
  • We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs.
  • Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation.

Sources (1)

  • [1]ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:39 AM
    To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks.
    Large language model agents are increasingly deployed to perform complex tasks in real-world environments.

Extractive summary: sentences quoted from the sources.