ProductsProduct / feature launchEvaluation & Benchmarks · MLOps, Tooling & Infrastructure1 source · Sep 30, 2026

Claude’s new auto eval tool

Anthropic released new eval tooling for Claude Code.

Proof1 independent outlet

Key points

  • Their claude-api plugin now includes a new buildeval and hill-climb command that helps you build evals, check the graders, and improve your application against them.
  • I usually don’t review eval tools because software changes so often that reviews have a short shelf life.
  • But a first-party tool from Anthropic is likely to influence how people approach evals, so I wanted to try it and share what I found.
  • In the screenshot below, Claude asks us to skim inputs.md and “tell it” which labels are wrong.

Sources (1)

  • [1]Claude’s new auto eval tool
    Hamel Husain · Sep 30, 07:00 AM
    Anthropic released new eval tooling for Claude Code.
    Their claude-api plugin now includes a new build_eval and hill-climb command that helps you build evals, check the graders, and improve your application against them.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 29, 2026langchain-ai/langchain langchain-anthropic==1.7.5
  2. Sep 29, 2026Sonnet 5.5 is worth a try
  3. Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
  4. Sep 29, 2026Claude Code’s Next Era — Thariq Shihipar, Anthropic
  5. Sep 29, 2026What do you want from AI?
  6. Aug 21, 2026ollama/ollama v0.33.0

Related