LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents.
Key points
- Evaluating agents by outcomes alone can obscure the capabilities that produce them.
- This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions.
- Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use.
- LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability
Sources (1)
- [1]LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving MarketsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:33 AM
We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents.
Evaluating agents by outcomes alone can obscure the capabilities that produce them.
Extractive summary: sentences quoted from the sources.