There’s an awkward problem with “improving” an AI agent: sometimes you make the benchmark better without making the agent better.
Change the prompt. Swap the model. Add a tool. Crank up the reasoning effort. Run the test again.
Score goes up.
Ship it?
Maybe.
Anthropic just published a much more disciplined way to answer that question. Its new claude-api eval and hillclimbing workflow gives Claude Code two jobs: first, help you build an evaluation that actually resembles the work your AI does in production; second, repeatedly modify your system while checking whether each “improvement” survives on examples it was never allowed to study.
Claude Developers shared the workflow here, and Anthropic’s Lance Martin, who wrote the underlying guide, walked through the method here.
The practical idea is almost painfully simple:
Real tasks → reliable grader → baseline → one change → test it → keep or revert.
Which sounds suspiciously like “do science,” except Claude is now also the lab assistant.
And Anthropic’s own examples suggest the payoff can be substantial. In one customer-support evaluation, the final configuration became more accurate while costing roughly one-fifth as much per ticket. In another, Claude iteratively improved Anthropic’s own API skill from about 66% to 88%.
The interesting part isn’t that Claude can tweak a prompt.
It’s the machinery Anthropic built around Claude to tell the difference between a useful improvement and a really convincing lie.
- First, what is an eval?
- The sneakiest eval mistake: testing whatever your model happens to be bad at
- /claude-api build-eval tries to automate the boring part
- The grader matters as much as the agent
- Then Claude measures the noise before trying to improve anything
- /claude-api hillclimb turns optimization into a controlled loop
- One change at a time
- And when improvement stalls, Claude stops randomly poking things
- The customer-support example is kind of wild
- Claude also hillclimbed Anthropic’s own Claude skill
- This starts to look like CI for AI behavior
- But automated hillclimbing does not solve the hardest problem
- And a better benchmark still isn’t the same thing as a better business
- The practical playbook
- The bottleneck is moving from prompts to measurement
First, what is an eval?
An eval is basically a test suite for AI behavior.
With normal software, the question might be straightforward: did this function return the correct number? Did the test pass?
Agents are messier. They write emails, investigate incidents, route support requests, browse websites, create reports, use tools, and make judgment calls. Two outputs can look different while both being correct. The same model can also answer the same task differently across multiple runs.
So before optimizing anything, Anthropic says your eval needs four properties.
- The tasks should look like production. If customers ask your support agent about refunds, damaged products, billing problems, and account access, your eval should test those things rather than a convenient pile of toy questions.
- Better models and more reasoning should generally score better. If your expensive frontier model performs worse than a smaller model, something may be wrong with the task, grader, configuration, or environment.
- There needs to be room to improve. If your strongest setup already scores around 95% or higher, measuring further quality gains gets difficult. Anthropic’s tooling warns you at roughly that point and suggests optimizing something like cost or latency instead.
- The score should be reasonably stable between runs. If the exact same system swings wildly between 65% and 90%, you may be measuring grader noise, ambiguous tasks, leftover environment state, or inconsistent model settings rather than a real product change.
That last point becomes especially important once you start letting an AI optimize the AI.
If your ruler changes length every time you measure something, an automated optimization loop can spend a lot of time getting extremely good at moving the ruler.
The sneakiest eval mistake: testing whatever your model happens to be bad at
Anthropic calls out a failure mode I think is especially easy to fall into.
Imagine you run your agent 1,000 times, collect the 50 ugliest failures, and turn those failures into your benchmark.
Sounds smart. Those are exactly the mistakes you want to fix.
But modern models have what Anthropic calls a jagged capability surface. They can be strangely brilliant at one task and strangely terrible at a nearly identical one.
If you select every eval case because the current model failed it, you may accidentally create a benchmark measuring the quirks of that particular model rather than the genuinely difficult parts of your product.
Then the next model arrives and your carefully curated benchmark becomes a museum exhibit.
Anthropic’s recommendation is to start from cases that humans can explain are genuinely difficult or valuable. Real production failures still belong in the eval, but you want a distribution that represents the job, not a scrapbook of whichever potholes Claude happened to hit last Tuesday.
There’s another trap: production traffic itself can skew too easy.
Users learn what a product can do. Then they stop asking it to do things they expect will fail.
So “sample real traffic” is useful, but even real traffic needs judgment.
/claude-api build-eval tries to automate the boring part
Anthropic’s first new workflow is /claude-api build-eval.
Run it inside Claude Code and Claude interviews you about what you are evaluating, builds the evaluation inside your codebase, and pauses at key moments for approval.
When assembling test cases, it prefers evidence in roughly this order:
- Production transcripts, after checking how sensitive data and retention should be handled.
- Bug reports and support tickets.
- Five to ten examples you write manually.
- Synthetic cases generated from the codebase.
Claude then generates a simple local page where you can inspect the proposed inputs before approving them.
That human review step is doing real work.
The goal is not “Claude wrote 100 test cases, therefore we have an eval.” You are checking whether those cases actually represent the thing you care about.
Then comes grading.
The grader matters as much as the agent
If your task has a constrained answer, Anthropic prefers programmatic grading.
Think:
- Did the correct label come back?
- Is the JSON valid?
- Did the test suite pass?
- Does the required file exist?
- Is the answer an exact match?
That’s great because code can judge code very consistently.
Open-ended tasks need something different.
If you’re evaluating a research report or customer response, many different outputs could be good. Anthropic’s workflow can use a second language model as an LLM judge. That judge gets the input, output, and a rubric written as specific checkable claims rather than a fuzzy 1-to-5 rating.
If there’s an existing baseline, the judge can instead compare the baseline and candidate in randomized order without being told which is which.
And Anthropic explicitly recommends not using the model being evaluated as its own judge.
The student does not also get to grade the final.
Before trusting the grader, Claude scores a sample of cases and asks whether you would have judged them differently. Anthropic says misconfigured scoring is one of the most common ways evals go wrong.
It also runs the grader twice on the same output to see whether the verdict changes.
That is a tiny detail with huge implications. If the grading system itself flips between pass and fail, optimizing against it is basically optimizing against a slot machine.
Then Claude measures the noise before trying to improve anything
Once the cases and grader look reasonable, the workflow runs a baseline and reports the score with a confidence interval, meaning a statistical range representing the uncertainty around the measured result.
This matters because an eval score can move even when your product hasn’t meaningfully improved.
Suppose version A scores 81%.
Version B scores 82%.
Is B actually better?
Maybe. Or maybe that one-point difference is normal randomness.
Before hillclimbing begins, Anthropic says Claude checks whether the eval’s noise is smaller than the smallest improvement you would actually care about. If the test is too noisy to detect that difference, the workflow recommends adding more examples or repetitions instead of pretending it can see signal that isn’t there.
That may be the least glamorous feature in the entire system.
It may also be the most important.
/claude-api hillclimb turns optimization into a controlled loop
Once you have an eval you trust, /claude-api hillclimb lets Claude start changing the application.
You decide what it is allowed to modify. Anthropic lists surfaces including:
- System prompts.
- Skills and instruction files.
- Tool descriptions.
- Model choice.
- Reasoning effort.
- Other API parameters.
- The surrounding agent harness, meaning the code and loop that connect the model to its tools and environment.
You also tell Claude what you actually care about.
Maybe you want maximum accuracy.
Maybe accuracy is already good enough and you want to cut cost.
Maybe you want similar performance from a cheaper model.
Then Claude randomly splits the evaluation into a train set and a held-out test set.
This distinction is the entire game.
Claude is allowed to inspect failures from the training examples.
It does not get to study the held-out answers while deciding what to change.
One change at a time
Each hillclimbing round follows roughly the same loop:
Inspect train failures → find a root cause → propose one patch → rerun eval → compare train and test → keep or revert.
Anthropic deliberately uses one change per round so you have some chance of knowing what caused the score to move.
If both train and held-out performance improve, the patch can stay.
If performance regresses, Claude reverts it.
And if the training score goes up while the held-out score stays flat, Claude treats that as a warning that it may be overfitting and reverts the change.
Overfitting here means your system learned the exam rather than the job.
Imagine one benchmark happens to contain several images that need OCR. The optimizer adds an OCR tool. Benchmark goes up.
Great.
Except your real customers almost never need OCR.
Or maybe several benchmark tasks happen inside /app, so the optimizer inserts a special instruction saying “always change into /app and run pytest.”
Again, benchmark goes up.
Production barely changes.
In the worst case, evaluation answers accidentally become accessible to the agent and the model simply finds them.
Anthropic explicitly warns about these kinds of leaks and tells the hillclimber not to paste individual failure content directly into the production prompt.
And when improvement stalls, Claude stops randomly poking things
After two or three rounds without progress, the workflow changes tactics.
Instead of immediately making another edit, Claude reads the remaining training failures and groups them by root cause.
Maybe several failures come from one missing rule.
Maybe an eval case itself is ambiguous.
Maybe the harness is broken.
Maybe the grader is wrong.
Maybe the difference you’re chasing is smaller than the eval can reliably measure.
Only failures that still look legitimate get fed into more optimization rounds.
When the process ends, Claude leaves the code at the version that performed best on the held-out set for the goal you chose.
Then it compares that version with the original baseline, including confidence intervals.
If the apparent gain still sits inside the eval’s noise, Anthropic says the workflow recommends not merging it.
Imagine an AI coding tool proudly finishing with: “I changed nothing. That was the correct decision.” Progress.
The customer-support example is kind of wild
Anthropic tested this on an internal customer-support benchmark containing 44 tickets.
Thirty tickets were used during the hillclimbing search. Fourteen were held out and never shown to the optimizer.
The starting setup used Opus 4.8 at high effort.
On the 30 search tickets:
- Accuracy: 74.4%
- Token cost: 4.6 cents per ticket
Claude first audited the prompt and removed things like mandatory tool-call rituals, a scratchpad step, and contradictory rules.
Then it tried Opus 5.5 at low effort.
That reached 87.8% accuracy for about 1.9 cents per ticket.
Once Opus 5.5 cleared the baseline, Claude tested whether it could step down to a cheaper model.
Sonnet 5 at low effort reached 88.9%, at roughly 1 cent per ticket.
Then Claude improved the routing prompt and added a refund-cap cross-reference.
Final training/search accuracy: 98.9%.
Still roughly one cent per ticket.
That result alone would be easy to over-celebrate.
Remember, Claude had seen those 30 tasks.
So Anthropic checked the 14 tickets it had held out.
The original setup scored 78.6% there.
The final setup scored 90.5%.
And it did so at about one-fifth the original cost.
That’s the result I’d write down.
Not 98.9%.
90.5% on the cases the optimizer never got to study, while spending dramatically less.
Claude also hillclimbed Anthropic’s own Claude skill
The second example may be even more instructive because the optimizer eventually discovered problems with both the product and the test.
Anthropic evaluated its claude-api skill, which teaches Claude how to use Anthropic’s APIs correctly.
Baseline performance was about 66%.
Claude found eight API features that the skill was missing. Adding them pushed performance to 74%.
Then it found errors in the C# and Java type tables.
Fixing those moved the score to 77%.
After progress stalled, Claude analyzed the remaining failures as a group.
It noticed that some instructions were technically present, but Claude still tended to write older API patterns from its pretraining. So the hillclimber added a table near the top explicitly mapping older remembered patterns to the current API.
That lifted performance to 80%.
Then came the fun part.
Some stubborn “failures” turned out to be bad evaluations.
One task asked Claude to catch one error type, while the grader expected a chain containing at least three.
Another grader contradicted Anthropic’s own documentation. Testing the real API showed the docs were right.
Fixing those evaluation problems, along with further skill changes, eventually brought the result to roughly 88%, with Anthropic’s chart showing 87.9% at round 24.
That’s a useful reminder of what this system is really optimizing.
Sometimes the agent is wrong.
Sometimes the instructions are wrong.
Sometimes the benchmark is wrong.
A serious eval workflow needs to be able to discover all three.
This starts to look like CI for AI behavior
Traditional software teams already have a version of this discipline.
Change the code.
Run the tests.
If something breaks, don’t merge it.
AI development has struggled to get the same feedback loop because the thing you are testing is probabilistic and often open-ended.
Anthropic is trying to make that loop much more mechanical:
Change the model system → evaluate real tasks → quantify uncertainty → test unseen cases → automatically revert regressions.
That matters more as agents get longer-lived and more complicated.
In our deep dive on Claude Managed Agents, teams from Wispr, Actively, and Pendo described the same basic transition. Early agent development can be vibes-based because you are still figuring out whether anyone wants the product. Once customers depend on the workflow, every prompt, model, memory, tool, and harness change becomes a possible regression.
That’s when evals stop being a research accessory.
They become part of the product infrastructure.
But automated hillclimbing does not solve the hardest problem
There’s one obvious catch.
Claude can only optimize the objective you give it.
If your eval rewards the wrong behavior, automated optimization can simply make your system better at the wrong behavior.
If your benchmark overrepresents easy cases, you’ll optimize easy cases.
If your LLM judge misunderstands the rubric, you’ll optimize toward its misunderstanding.
If your real product changes every day because the agent depends on live Slack messages, changing memories, customer data, or external websites, an offline held-out set may still fail to represent reality.
The teams in our Managed Agents piece described exactly this problem. Actively said traditional offline evals become awkward when agent answers depend on live memory, and Wispr faces similar issues when connected services change underneath the test.
Anthropic’s workflow reduces one class of self-deception.
It does not make measurement magically objective.
And a better benchmark still isn’t the same thing as a better business
This distinction gets lost constantly in AI.
Model score improved.
Great.
Did customers complete more work?
Did humans accept the output?
Did failures become cheaper?
Did the workflow become reliable enough to run unattended?
Did the economics actually get better?
We went deeper on that distinction in our guide to measuring AI ROI without fooling yourself. A model can look excellent at an individual step while a long workflow remains fragile, because errors compound across every step the system has to complete.
Anthropic’s support example is compelling precisely because it optimizes two things you can actually care about in production:
accuracy and cost.
That’s much more useful than moving a benchmark from 92 to 93 because 93 looks nicer in a launch deck.
The practical playbook
If I were using Anthropic’s approach on a real agent tomorrow, I’d steal the workflow more than the exact commands:
- Start with the actual job. Pull examples from production traces, support tickets, bugs, and tasks humans genuinely care about.
- Make the grader boring. Use code whenever possible. Use an LLM judge only when the task genuinely requires judgment, and inspect its decisions manually before trusting it.
- Check the eval before optimizing the agent. Stronger models should usually do better, variance should be manageable, and there should be enough headroom to detect improvement.
- Measure your noise floor. Decide the smallest improvement worth acting on, then make sure your test can actually distinguish that change from randomness.
- Choose one surface to optimize. Prompt, skill, model, effort setting, tool description, or harness. Smaller reversible surfaces are much easier to reason about.
- Split what Claude can see from what proves the result. Let it inspect training failures. Keep the held-out examples hidden.
- Change one thing at a time. Run the eval after each patch.
- Revert aggressively. Train up and held-out flat? Probably overfitting. Test down? Revert. Difference inside the noise? Don’t pretend.
- When progress stalls, inspect the failures themselves. You may have found an agent problem, an instruction problem, a grader problem, or a benchmark problem.
- Keep one metric tied to reality. Cost per completed task, accepted output, resolved ticket, successful workflow, human intervention rate, whatever your business actually cares about.
That last one keeps the whole optimization loop attached to the ground.
The bottleneck is moving from prompts to measurement
For the last few years, a huge amount of practical AI advice has boiled down to prompt craft.
Use clearer instructions.
Add examples.
Give the model more context.
Pick a better model.
All still useful.
But once agents become capable enough, the problem changes.
You can make a hundred plausible improvements.
The hard part is knowing which one actually worked.
Anthropic’s build-eval and hillclimb workflow is an attempt to automate that discipline. Claude can generate the tests, help validate the grader, measure the baseline, alter the system, inspect failures, compare train and held-out performance, revert regressions, and tell you when the result is too noisy to trust.
That makes agent development look less like prompt engineering and more like running an experiment.
Which leaves one thing humans very much still own:
deciding what deserves to count as success.
Because if the eval lies, Claude can now optimize the lie much faster.