Claude’s new auto eval tool
Anthropic released new eval tooling for Claude Code.

Proof1 independent outlet
Key points
- Their claude-api plugin now includes a new buildeval and hill-climb command that helps you build evals, check the graders, and improve your application against them.
- I usually don’t review eval tools because software changes so often that reviews have a short shelf life.
- But a first-party tool from Anthropic is likely to influence how people approach evals, so I wanted to try it and share what I found.
- In the screenshot below, Claude asks us to skim inputs.md and “tell it” which labels are wrong.
Sources (1)
- [1]Claude’s new auto eval toolHamel Husain · Sep 30, 07:00 AM
Anthropic released new eval tooling for Claude Code.
Their claude-api plugin now includes a new build_eval and hill-climb command that helps you build evals, check the graders, and improve your application against them.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 29, 2026langchain-ai/langchain langchain-anthropic==1.7.5
- Sep 29, 2026Sonnet 5.5 is worth a try
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Claude Code’s Next Era — Thariq Shihipar, Anthropic
- Sep 29, 2026What do you want from AI?
- Aug 21, 2026ollama/ollama v0.33.0