Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds
To address these challenges, we introduce E-Ledger, a multi-agent harness for safe and persistent execution.
Key points
- Large language model agents can invoke tools fluently, but enterprise workflows demand more than selecting the right tools: actions must strictly comply with organizational policies, tool feedback often conceals hidden side effects under partial observability, and long-horizon tasks require persistent state tracking across multiple records.
- Because hidden rules are typically unknown a priori, we further propose WorldAbduct, an abductive, world-model-driven harness evolution framework.
- WorldAbduct diagnoses execution trajectories across four complementary views (state consistency, world-observation gap, policy-gate correctness, and goal judgment) to hypothesize latent rules, and verifies them through targeted abductive interactions before integrating them into the ledger.
- On the enterprise benchmark World of Workflows, E-Ledger with WorldAbduct improves safe task completion across four LLM backbones, outperforming the strongest evolution baseline by 5--15 percentage points.
Sources (1)
- [1]Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable WorldsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:15 AM
To address these challenges, we introduce E-Ledger, a multi-agent harness for safe and persistent execution.
Large language model agents can invoke tools fluently, but enterprise workflows demand more than selecting the right tools: actions must strictly comply with organizational policies, tool feedback often conceals hidden side effects under partial observability, and long-horizon tasks require persistent state tracking across multiple records.
Extractive summary: sentences quoted from the sources.