Humanize: Judgement Engineering for Agentic Coding
We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning.
ProofPaper ↗
Key points
- Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done.
- A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work between these roles and enforce 72 mechanical gates.
- Viewed as a Markov chain over repository states, alternating builder and reviewer samples jointly from two models, so a defect survives only if both miss it.
- We study Humanize through its deployment, 118 public postmortems of real loops, and its applications.
Sources (1)
- [1]Humanize: Judgement Engineering for Agentic CodingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:54 PM
We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning.
Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done.
Extractive summary: sentences quoted from the sources.