One Step at a Time: Trading LLM Autonomy for Process Predictability
Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step.
Key points
- When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back.
- We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured stepoutput.
- Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small.
- Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.
Sources (1)
- [1]One Step at a Time: Trading LLM Autonomy for Process PredictabilityarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:10 AM
Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step.
When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back.
Extractive summary: sentences quoted from the sources.