AION
Research paperAgents & Tool Use · Safety & Alignment · Reasoning & Planning1 source · Oct 7, 2026

LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents.

Key points

  • Evaluating agents by outcomes alone can obscure the capabilities that produce them.
  • This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions.
  • Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use.
  • LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability

Sources (1)

Extractive summary: sentences quoted from the sources.