Agentic evals: how to know whether an AI agent did the job, and will do it again

@sermakarevich
Sergii Makarevych@sermakarevich
24 views Oct 10, 2026 ~44 min read
Advertisement

An article for engineers who ship agents built on large language models (LLMs): systems that call tools, change records, and act on a customer's behalf over several steps. After reading it you will be able to grade an agent run by what it changed rather than what it said, put a reliability number next to the capability number, find the step where a failing run went wrong, and tell when a model-as-judge can be trusted. It then widens to their managers and to whoever decides whether an agent may touch orders, code, or money.

Media image

The numbers come from published papers listed at the end and from one runnable experiment on a small support agent that Level 0 describes completely. Nothing here assumes you have seen it before.

The quick grasp (read this if you read nothing else)

An agent is a language model in a loop: it reads a request, picks a tool (look up an order, cancel it, refund it, send a reply), reads the tool's answer, and picks again until it decides it is done. The code around the model that runs this loop, with its prompt, tools and retries, is the harness. An agent eval is a repeatable test that tells you, with a number you can trust, whether the agent did the job, whether it does it every time, and where it goes wrong when it does not.

The previous article, Evals: how to know whether an AI system actually works, covered systems that give one answer to one question: how to build a test set, grade an answer, trust a model-as-judge, and read the numbers. All of that still applies here. What changes is what can go wrong.

A single-answer model can only be wrong in its text. An agent can be wrong in its text, its actions, their order, their count, and in what it left behind in the database, and these disagree. The reply can say "your refund is on its way" while no refund exists. In the example used throughout this article, a grader that only reads the reply passes 58 of 60 runs; a grader that checks the database passes 48. The other 10 sounded right while the database said otherwise.

How it works, in five steps:

  • Write tasks with an expected end state. Not "the customer feels helped" but "order O101 is cancelled, 120 dollars refunded, no fee charged".
  • Run the agent on each task several times from the same starting database, and keep the whole trail: tool calls and results, final reply, database before and after.
  • Grade the end state and the path. Did the records change the right way? Were the required tools used and the forbidden ones avoided? Did it finish inside the allowed number of steps? Only then read the reply.
  • Report two numbers. How many tasks it solves at least once (capability), and how many it solves every time (reliability). The second is the one customers experience.
  • For each failure, find the first wrong step, count the kinds, and fix the most common kind first.
  • When you use it: before the agent acts for real, before every change to prompt, tools, or model, and continuously on sampled real traffic once it is live.

    Important: grade what the agent changed, not what it said. The reply is the last and least reliable piece of evidence.
    What it means for us: every task needs a checkable end state and a list of allowed actions, written before the agent runs. A task without them is not yet an eval.
    Why it matters: a do-nothing agent that only writes a friendly reply passes every reply-only check in the example, 20 tasks of 20. A number that cannot fail a do-nothing agent is not a number.

    The rest of the article zooms in one level at a time.

  • Level 0 · One task, end to end: the shop, the tools, the grader, the numbers
  • Level 1 · What counts as "did the job"?
  • Level 2 · Will it do it again?
  • Level 3 · Where did it go wrong?
  • Level 4 · Can a model grade an agent?
  • Level 5 · What did each success cost?
  • Level 6 · What changes when agents hand off to agents?
  • Level 7 · What if the agent learns to beat the grader?
  • Level 8 · Making it run every day
  • Then · The tools, the public benchmarks and how to read them, a decision guide, and the references
  • Level 0: one task, end to end

    This level answers: what does an agent eval look like when it runs, from the customer's message to the two headline numbers?

    The example is a customer-support agent for a fictional bike shop, Rivertown Cycles. It is small enough to describe completely and realistic enough to fail the way real agents fail.

    The world. A mock database of about twenty customer orders, each with a status (placed, processing, shipped, delivered, cancelled, returned), a total, an address, items, a refund total, and a return status. The database is rebuilt before every run, so every run starts from the same world.

    The tools. The agent has 8 tools, and each one enforces shop policy inside itself:

  • lookup_order and list_orders read the database and change nothing.
  • cancel_order works only while an order is placed or processing, with a 5-dollar fee once processing has started.
  • update_address works only while an order is placed, within 24 hours of placing it.
  • start_return works only on delivered orders, within 30 days, never on custom-built items, and moves no money.
  • issue_refund is capped at the order total, cumulative across refunds.
  • escalate hands the case to a human supervisor and returns a ticket id.
  • reply sends the final text to the customer and ends the run.
  • Policy inside the tools means the agent can ask for a change the rules do not allow, the tool refuses, the database stays clean, and the refusal is logged as a tool error.

    The instructions. One system prompt gives the date, the tool order (look up the order if an id is mentioned, otherwise list the customer's orders), the rules above, one more rule that any refund of 200 dollars or more not tied to a started return must be escalated first, and that the run must end with one reply that invents nothing.

    The model. A downloadable mid-sized model (about 27 billion parameters) on one local graphics card, with enough built-in randomness that repeated runs can differ. A bigger model moves the numbers, not the lessons.

    The tasks. 20 hand-written customer messages, each with a gold specification written before any run: the expected end state as database fields and values; required tools; forbidden tools; a step budget of 8 to 12; and a few keywords the reply must contain. The tasks fall into 11 straightforward actions (cancel, change address, return, look up), 7 policy refusals (change nothing and explain why), and 2 escalations (go through the supervisor).

    The run. Every task runs 3 times from the clean world, which gives 60 episodes. An episode is one full loop from customer message to final reply. The whole trail is stored: each tool call with arguments and result, the model's reasoning when it wrote any, the final reply, the database before and after, and their difference.

    One episode. Task t01: "Hi, I just realized I never paid for the helmet in my order O101, and I want it cancelled. Please cancel it." The expected state is status cancelled, refund total 120, fee 0. The run goes:

  • lookup_order(O101). Result: status placed, placed 1.1 days ago, total 120, refund total 0.
  • cancel_order(O101). Result: status cancelled, fee 0, refund 120.
  • reply("Hi Ana, I've cancelled order O101 (Trail helmet + Gloves, $120.00). No cancellation fee applied since the order was still in placed status…").
  • The grader. Pure code. Five checks, and the episode passes only if all five do: the expected state is in the database after the run; every required tool appears in the trail; no forbidden tool does; the step count is inside the budget; the reply exists and contains the keyword. The episode above passes all five in 3 of 8 steps.

    The numbers. Over the 60 episodes:

  • 48 of 60 episodes pass all five checks: a pass rate of 0.80, with an interval from 0.62 to 0.95. The interval is the range the rate would likely land in with a different 20 tasks of the same kind; it is wide because 20 tasks is a small set. It is computed by redrawing tasks, not episodes, because the 3 episodes of a task rise and fall together.
  • 17 of 20 tasks pass at least once in 3 tries (0.85). This is capability.
  • 15 of 20 tasks pass all 3 tries (0.75). This is reliability.
  • 2.5 tool calls per episode on average, maximum 5, far inside the budget.
  • In 1 of 60 episodes the agent asked a tool for a change the rules do not allow, a refund above the order total; the tool refused and the agent corrected itself. That is the "helpful agent makes a wrong change" case the design exists to catch.
  • Carry forward from Level 0:

  • An episode has five things to grade: end state, required tools, forbidden tools, step count, reply. Grade them in that order.
  • Start every run from the same world and keep the whole trail, or you cannot grade the path or debug a failure later.
  • Two headline numbers: solved at least once (0.85) and solved every time (0.75). Report both.
  • Level 1: what counts as "did the job"?

    This level answers: which evidence should the grader look at, and what happens when it looks at the wrong one?

    The five checks in Level 0 are not equally strong. The example makes this measurable: the same 60 episodes were re-graded four ways, and a second "agent" was added that does nothing at all.

    Four graders over the same 60 episodes.

  • The full grader (all five checks) passes 48 of 60.
  • A reply-only grader, which looks for the expected keywords in the final text, passes 58 of 60. Ten episodes that sound right are counted as done.
  • A state-only grader, which checks only the database fields, passes 48 of 60, the same 48: every failure here left the database wrong.
  • A model-as-judge reading the transcript passes 49 of 60, but a different 49: it passes 12 episodes the full grader fails and fails 11 the full grader passes.
  • Where the reply-only grader goes wrong. All 10 wrongly passed episodes fail the state check, and 8 also skipped a required tool. The pattern repeats: the agent looks up the order, writes a warm reply that promises a refund or confirms a return, and never calls the tool that would do it.

    The do-nothing agent. To test a grader, feed it an agent that cannot be doing the job. A null agent that calls no tool and replies with the expected keyword was graded the same ways over the 20 tasks:

  • Reply-only grader: 20 of 20 pass.
  • State-only grader: 11 of 20 pass: all 7 policy refusals, because "change nothing" is the correct end state there, and 4 read-only tasks whose expected state is the starting state.
  • Full grader: 6 of 20 pass, the six refusal tasks that list no required tool. The seventh refusal requires a look-up and so fails.
  • The second line is the one to remember. State grading is necessary, but alone it cannot tell "refused for the right reason" from "did nothing"; 11 of the 20 tasks, 33 of the 60 episodes, have an expected state equal to the starting state. On those, required tools and the reply check carry the weight, which is why the full grader needs all five.

    The same pattern in the field.

  • The ABC audit (Agentic Benchmark Checklist, Zhu et al., 2025) found a trivial agent passing 38% of the airline tasks in τ-bench, a public customer-service benchmark, for the same reason the null agent passes refusals above. Its checklist asks every benchmark two questions: can the task be solved (task validity), and does the grader detect solutions (outcome validity). Of 10 benchmarks audited, 7 had an outcome-validity flaw.
  • AgentLens (2026) re-graded 2,614 agent trails by hand and found 10.7% of "passes" were lucky: right end state, wrong path, with three annotators in near-perfect agreement. An end-state grader counts a lucky pass as a pass.
  • What the full grader still misses. Tone, whether a correct action was explained correctly, harmless detours: jobs for a judge with the right evidence and for trail metrics.

    What it depends on. The gold specification being right, and the example has two tasks where it is not. Task t14 asks for a return on order O104 to be "confirmed in the system" and expects the return to exist, but forbids starting one; on a fresh database no return exists, so no agent can pass it. Task t17 expects a 240-dollar refund on order O110 that only an earlier task would have created. Both were written as if the previous task's changes persisted, and they do not.

    Of the five tasks that fail or flake in Level 0, three are agent failures and two are task bugs. Without those two tasks the pass rate rises from 0.80 to 0.89 and reliability from 0.75 to 0.83. The ABC checklist's first question, applied before any run, would have caught it; the null agent does not, because it fails those tasks too.

    Important: a grader is only as good as the worst shortcut it takes. Reply-only grading passes a do-nothing agent on every task; state-only grading passes it on half.
    What it means for us: build the grader from state, required actions, forbidden actions, and budget first, with the reply as the last check. Then run a do-nothing agent through it, and check every task against the clean world, before trusting any number it produces.
    Why it matters: the two most common ways to fool yourself with an agent are a fluent reply and an unsolvable task. One is caught by the null agent, the other by the clean-world check.

    Carry forward from Level 1:

  • Grade end state and path. A reply-only grader cannot fail a do-nothing agent.
  • State alone is blind on "change nothing" tasks, 11 of 20 here. Required actions and a reply check cover them.
  • Run a null agent through every grader and check every task is solvable from the clean world. Two of 20 tasks here were not.
  • Level 2: will it do it again?

    This level answers: how do you turn repeated runs into a capability number and a reliability number, and how many runs do you need?

    Level 0 reported 0.85 and 0.75 from the same 60 episodes. They come from one estimator each.

    pass@k and pass^k. Run each task n times and count the passes c.

  • pass@k is the chance that at least one of k fresh tries passes. It measures capability: the agent can do this.
  • pass^k (read "pass to the k") is the chance that all k fresh tries pass. It measures reliability: the agent does this every time. With n trials and c passes it is the share of k-sized sets of trials that are all passes, averaged over tasks. A task with 2 passes in 3 trials has 3 possible pairs and 1 all-passing pair, so its pass^2 is one third. That is τ-bench's estimator, exact for any k up to n.
  • With 3 trials per task the example gives:

  • k of 1: 0.80 for both, interval 0.62 to 0.95. One try is one try.
  • k of 2: pass@2 0.83, pass^2 0.77.
  • k of 3: pass@3 0.85 (interval 0.70 to 1.00), pass^3 0.75 (interval 0.55 to 0.90).
  • The gap between the lines is the reliability gap. It is zero at k of 1 by construction and opens as k grows, pass@k climbing toward "tasks the agent ever solves" and pass^k falling toward "tasks it never fails". With 3 trials it has only two steps to open in; with 5 or 8 the two lines separate further.

    Per task, the gap has names. 15 tasks pass all three trials; 2 are flaky (t04, a return plus refund, passes 1 of 3; t10, an escalation plus refund, 2 of 3); 3 never pass (t09, and the two task bugs t14 and t17). Per category:

  • Policy refusals, 7 tasks: pass@3 1.00, pass^3 1.00. Refusing is reliable.
  • Straightforward actions, 11 tasks: pass@3 0.82, pass^3 0.73.
  • Escalations, 2 tasks: pass@3 0.50, pass^3 0.00. The agent escalates t10 in 2 of 3 tries; t17 is one of the two tasks no agent can pass.
  • The headline 0.75 is a weighted average of 1.00, 0.73 and 0.00. If the escalation path is the one that moves large refunds, the headline is not the number to ship on.

    Why 3 trials is the floor, not the recommendation. With k equal to n, pass^k is all or nothing per task: three lucky passes make a flaky task look reliable. Five trials smooth pass^3 and make pass^5 computable. The intervals above redraw tasks, not episodes, for the reason given in Level 0.

    Paraphrase the request. A second kind of repetition changes the input instead of the random seed. The same 20 tasks are rewritten with the same intent in different words ("I want it cancelled" becomes "please stop this order going out, I changed my mind") and run again against the same specifications. For the example this run is in progress as this is written, so no number is reported here. One phrasing per task measures the agent against your words, not the customer's.

    The same gap in the field.

  • Rabanser and colleagues (2026) measured 15 models over 5 runs each: accuracy improved six times faster per year than consistency.
  • τ-bench (2024) reports GPT-4o at pass^1 of 61.2 on retail and 35.2 on airline, and under 25 for pass^8 on both. An agent that succeeds six times in ten succeeds eight times in a row less than one time in four.
  • A study of code-generation robustness (2023) found 46% of meaning-preserving paraphrases of a prompt changed the generated code. The seed is not the only source of variance; the wording is.
  • Important: capability and reliability are different numbers from the same runs, and the second is what the customer meets.
    What it means for us: run every task at least 3 times, 5 if you can, report pass@k and pass^k with task-clustered intervals, and break them out by category. A single-run pass rate is a capability number in disguise.
    Why it matters: the example's 0.85 and 0.75 hide a 1.00 on refusals and a 0.00 on escalations. The average is true and useless.

    Carry forward from Level 2:

  • pass@k measures "can it", pass^k measures "does it, every time". The gap grows with k.
  • Three trials is the floor; five lets you see pass^5 and smooths pass^3. Resample tasks, not episodes.
  • Vary the wording as well as the seed. One phrasing per task measures the agent against your phrasing, not the customer's.
  • Level 3: where did it go wrong?

    This level answers: once an episode fails, how do you find the step that broke it, and what do you count across failures?

    A pass rate says how often. The trail says where. Each of the example's 12 failing episodes can be walked back to its first wrong step: the earliest tool call not seen in a passing run of the same task, or a reply that came too early.

    The 12 failures, by first wrong step.

  • 6 are the two task bugs from Level 1: the three runs of t14 and the three of t17. On t14 the agent did the right thing and the specification was wrong. On t17 the agent also skipped the required escalation, but no behaviour could have reached the expected state.
  • 6 are real failures, and all six skipped issue_refund and went straight to the reply at step 3. Two of three runs of t04 (return a saddle, refund 180) and all three of t09 (return shoes, refund 140) looked up the order, started the return and replied. One run of t10 (a 240-dollar refund that policy says to escalate) started a return instead, then told the customer that "policy issues the refund only after we receive and inspect the returned item". No such policy exists. The agent invented a rule to fit the tool it chose.
  • 0 called a forbidden tool, and none of the 6 real failures asked a tool for a change the rules refuse. The failures are omissions, not transgressions.
  • That is the fix. One kind of mistake, "stops before the money moves", explains all 6 real failures and points at the prompt's refund rules, not at the model's tool-calling. Without the trail, the same 12 failures are a pass rate of 0.80 and a guess.

    Trail metrics to count on every run. Each is cheap, and each caught something here:

  • Steps per episode, mean 2.5 and maximum 5 against a budget of 8 to 12. A rising mean on the same tasks means looping.
  • Tool error rate, 7 of 151 calls: 6 were empty look-ups for a customer with no orders, correct on those tasks and a warning anywhere else, and 1 was the refused over-cap refund from Level 0.
  • Refused change attempts, counted separately from tool errors: the agent asked a tool for a change the rules forbid. One in 60 here; on policy tasks the gate is zero.
  • Reply-before-action: an episode ending in a reply with no state change on a task that expected one. That is the 6-episode pattern above, and it can be flagged automatically.
  • Failure taxonomies in the field. The categories differ by domain; the method is the same: label the first wrong step, count.

  • MAST (2025) labelled more than 1,600 multi-agent traces into 14 failure modes in 3 groups: specification, inter-agent misalignment, verification. The most common were step repetition (15.7%), reasoning that did not match the action taken (13.2%), and not noticing the task was done (12.4%): agents that loop, drift, or overrun, each visible only in the trail.
  • τ-bench's authors read 40 failures: 55% were a right tool with a wrong argument. That share is why Level 1 grades state, not tool names.
  • Model or Harness? (2026) sorted 41 failure modes, 36 on the model side. Its sharper finding: model judges asked to assign blame over-blame the model.
  • MP-Bench (2026) had 3 experts localise the root cause in 289 agent logs. They agreed on the step only 16.2% of the time, and the best model ranked the true cause in its top 5 at 0.44 against 0.13 for random. Finding the first wrong step is hard for humans too. It is still the only method that produces a fix.
  • Important: a failure is a step, not a score. Find the first wrong step in every failing episode and count the kinds.
    What it means for us: store the full trail by default, write one classifier for "replied without acting", and read the first wrong step of every failure by hand until the categories stabilise. Here one category explains every real failure.
    Why it matters: fixing the most common first wrong step is the cheapest improvement available, and no aggregate number can find it.

    Carry forward from Level 3:

  • Walk each failure back to its first wrong step and count the kinds. Here, every real failure is "replied before the money moved".
  • Count steps, tool errors, refused change attempts, and empty look-ups on every run. They move before the pass rate does.
  • Do not let a model assign blame unchecked. Judges over-blame the model, and experts agree on the root cause one time in six.
  • Level 4: can a model grade an agent?

    This level answers: when a second model judges an agent run, what evidence does it need, and how do you check it?

    The code grader is exact and blind: it cannot read tone or tell a good refusal from a curt one. A model as judge can. The question is whether it reads the right thing.

    The transcript judge in the example. A second model saw each episode's customer message, the numbered trail of tool calls with results, and the final reply, and returned pass or fail with a one-line reason. Over the 60 episodes:

  • The judge passes 49 and the code grader 48, so on the headline they agree.
  • They agree on only 37 of 60 episodes (61.7%). Cohen's kappa, agreement beyond chance, is −0.24: the two disagree slightly more than two coin flips would.
  • The judge passes all 12 episodes the code fails, the 6 task bugs and the 6 real misses alike. It reads a reply that promises a refund and credits the intent.
  • The judge fails 11 episodes the code passes: clean cancellations and correct refusals where it objected to the wording.
  • Lenient where it matters, strict where it does not. Two causes are visible. The judge never sees the database, so it cannot know the refund did not happen. And its instructions promised it "a summary of the database changes", which the pipeline never supplied. That mismatch was found by reading the prompt against the code, not by any metric. Fixing it is the first rung of a ladder: the judge's evidence grows from reply, to trail, to trail plus state difference, to all of that plus the gold specification.

    Adding the state to the evidence. The same 60 episodes were judged again by the same model, now given the database rows of every order the episode touched, before and after, with the list of changes.

  • Agreement with the code grader rose from 37 to 46 of 60, and wrong fails fell from 11 to 2. With the state in hand the judge stopped objecting to clean cancellations and correct refusals.
  • Kappa moved only from −0.24 to −0.06, because the judge still passed all 12 episodes the code failed.
  • Those 12 explain the ladder's third rung. Six are the runs of t14 and t17, where the judge is right and the specification is wrong: it saw no return on file and no eligible discount, and approved the honest reply. The other six are the real misses, where the agent started a return, promised a refund "once the item is received", and never issued one. The judge saw the missing refund in the difference and approved it anyway. State evidence removes the judge's false alarms; only the specification can show it the misses.

    The field agrees on the shape of the ladder.

  • GAUGE (2026) measured judges on simulated support conversations: 57.5% of conversations a judge rated as satisfying had failed the task, and judges from the agent's own model family rated it higher than outside judges did. Reply-reading judges grade the reply.
  • AgentRewardBench (2025), 1,302 web-agent trajectories with expert labels: of the passes the best judge declared, 69.8% were real, and it found 83.1% of the real passes; rule-based checks were right 83.8% of the time when they said pass but found only 55.9% of the passes. Every judge over-estimated success. Rules miss passes, judges invent them.
  • Agent-as-a-Judge (2024) gave the judge the agent's artefacts and intermediate outputs on 55 development tasks and reached about 90% alignment with human judgment against about 70% for a final-answer judge.
  • MAST's failure-mode judge reached a kappa of 0.77 against human labels with a few labelled examples in the prompt, and 0.58 without. Same model, same traces; the evidence in the prompt moved kappa by 0.19.
  • Three rules that follow. The code grader is the gate and the judge a second column, never a replacement. The judge receives the state difference and the specification, or it is a reply-only grader with extra steps. And the model that ran the episode never grades its own episode.

    Important: a judge is a grader with the evidence you gave it. Given the reply, it grades the reply.
    What it means for us: give every judge the trail, the state difference, and the gold specification; check it against the code grader on every run; publish the kappa. Under 0.6 the two are measuring different things, and you need to know which one you ship on.
    Why it matters: the example's judge matched the code grader on the headline, 49 against 48, and disagreed on 23 of 60 episodes. A headline match is not agreement.

    Carry forward from Level 4:

  • In the example a transcript-only judge reached a kappa of −0.24 with the code grader: it credits promised actions that never happened.
  • Evidence is the lever: with the state difference in hand the judge's wrong fails fell from 11 to 2; without the specification it still caught 0 of the 12 code failures.
  • Judges over-estimate success and over-blame the model; rules under-count passes. Report both, and never let the agent's own model judge it.
  • Level 5: what did each success cost?

    This level answers: how do you put a price on a passing run, and why is the average cost the wrong number?

    An agent that passes 80% of tasks at 2 tool calls each and one that passes 80% at 12 are not the same product. The example stores every model call, so cost can be attached to every episode and divided by the passes.

    Calls per success in the example. Each tool call is one model call, so the trail prices the episode. Token counts (the text pieces a model reads and writes) were not stored, so calls stand in for tokens here.

  • 151 calls over 60 episodes: 2.52 per episode, 3.15 per passing episode. The difference is the failed episodes' calls, charged to the successes, which is where they belong.
  • By category: 2.14 calls per pass on policy refusals, 3.56 on straightforward actions, 8.50 on escalations, where 17 calls bought 2 passes.
  • By task: from 2.0 calls per pass to 10.0 on t04, one pass in 3 tries. A 5-fold spread on tasks the specification rates as equally routine.
  • The two tasks no agent could pass have an infinite cost per success and an ordinary cost per episode, which is the point. Cost per episode hides the tasks that eat the budget.

    What the field has measured.

  • Bai and colleagues (2026) measured agent spend across 3,500 tasks: a 30-fold cost spread between models completing the same task, and a correlation of only 0.39 between a model's own estimate of its cost and the real cost. Agents cannot budget themselves.
  • The Harness Effect (2026), a study of 22 engineers, found a tuned harness around the same model cut cost 41% and time 44% while the pass rate moved from 0.78 to 0.81. A small sample, but it shows the loop around the model can move cost more than the model does.
  • The number to report. Cost per passing episode, by task category, with the maximum-to-minimum ratio across tasks, and seconds per episode next to it, because the customer waits for all of them. Average cost per episode rewards an agent that gives up early; cost per pass does not.

    Important: price the successes, not the episodes. An agent that quits cheaply looks efficient on the wrong metric.
    What it means for us: log tokens, model calls, and seconds per episode, divide by passes per category, and gate on the ratio between the most and least expensive task.
    Why it matters: a 30-fold spread in cost for the same task is a measured fact, and agents cannot predict their own spend.

    Carry forward from Level 5:

  • Cost per passing episode, by category, is the metric. Cost per episode rewards giving up.
  • Log tokens, calls, and seconds per episode. Latency is paid by the customer.
  • The loop around the model is a cost lever: the same model with a tuned loop cost 41% less in one small study.
  • Level 6: what changes when agents hand off to agents?

    This level answers: when one agent's output becomes another agent's input, where does the pass rate go, and whose fault is the loss?

    Most production agents are pipelines: a planner and a coder, a triage agent and a specialist. The question becomes "which agent lost it".

    A relay of the example. Split the support agent in two. A reader agent gets the customer message and the read-only tools and writes a short brief: who the customer is, which order, what they want, what policy says. An actor agent gets only that brief and the action tools and must finish the job. Same tasks, same grader, same clean world. This is the design; its run is queued and no result is reported here. What it adds is attribution: every failure now has three possible owners, and the stored brief is the evidence that decides between them.

    What the field measures when agents hand off.

  • A reliability-limits study of multi-agent planning (2026) passed the same task through relays of 2, 3 and 5 agents handing off in free prose, and a single agent's 90.7% became 41.2%, 43.5% and 22.5%. The authors put the loss at about 8.5 points per prose hand-off; a structured hand-off with fixed fields cost about 2.8 and kept a 3-agent relay at 75.2%. The format of the brief has a measured price.
  • The Specification Gap (2026) split coding tasks between a planner and an executor: the pair trailed a single agent by about 30 points at two difficulty levels (58.2 against 88.6 on the easier one). Giving the executor the full original specification instead of a summary brought the pair back to 88.9. Hand-offs lose the specification, and that loss is most of the gap.
  • An equal-token comparison (2026) found a single agent at 0.418 against a multi-agent system at 0.388 at the same token budget. Most multi-agent wins in the literature are also token wins.
  • A scaling study (2025) of 180 multi-agent configurations found a mean change of −3.5% against a single agent with a standard deviation of 45.2%: large gains and large losses averaging to nothing. Gains vanished once the single agent was above about 45%, and independent agents amplified errors 17.2-fold against 4.4-fold under a central coordinator.
  • Attribution, in practice. The relay gives three hypotheses for every failure, and the stored trail decides:

  • Hand-off loss: the brief is wrong or incomplete. Check: does it contain the order id, the request, and the policy outcome?
  • Actor error: the brief is right and the actor still failed. Check: the actor's first wrong step against the brief.
  • Grader error: both did the job and the specification is wrong. Check: the Level 1 audit.
  • Level 3's expert-agreement figure is a warning that this is hard even with the trail. Without the trail it is impossible.

    Important: every hand-off is a lossy channel. Measure the loss per hand-off, and give the receiving agent the original specification, not a summary of it.
    What it means for us: grade a pipeline end to end with the same state grader as a single agent, store every intermediate brief, and keep a single-agent baseline at equal token budget. Add an agent only when the pipeline beats that baseline on pass^k.
    Why it matters: in the field, prose hand-offs cost 8.5 points each and multi-agent layouts average −3.5% with a 45-point spread.

    Carry forward from Level 6:

  • Hand-offs lose specification. Structured briefs lose about a third as much as prose; sending the full original request loses least.
  • Keep a single-agent baseline at equal tokens. Multi-agent gains that survive it are rare and concentrate on tasks the single agent fails.
  • Attribute each failure to hand-off, actor, or grader from the stored trail, never from the agents' own account.
  • Level 7: what if the agent learns to beat the grader?

    This level answers: what happens when an agent is optimised against its own eval, and which defence holds?

    Everything above assumes the agent is trying to do the task. An agent tuned or trained to raise a number will, if the number can be raised without doing the task, do that instead. This is reward hacking, and it is measured.

  • The Verification Horizon (2026) studied coding agents trained against test-based rewards. With the tests visible and fixed, 28.57% of "resolved" tasks were hacked: the agent edited the tests, special-cased the inputs, or found the answer in an artefact. With verification kept out of the agent's reach and extended, the rate fell to 0.56%.
  • AEVO (2026), an automated harness optimiser, found its agent hacking the grader in 2 of 3 runs whenever there was no wall between the agent's workspace and the grading code.
  • The Meta-Agent Challenge (2026) had agents design agent harnesses; 5 trials were caught copying test answers into the harness. The authors kept a human in the loop on every design for that reason.
  • RRSI (2026) found a harness tuned on a benchmark scoring 92.8 on it and 40.3 on unseen tasks of the same kind. Tuning on the eval produces a number about the eval.
  • The one defence that holds. Where a defence was tested, in the Verification Horizon and AEVO studies, the hacking stopped when the thing the grader checks was out of the agent's reach. The example has that property by design:

  • Policy lives inside the tools, so the agent cannot talk its way past a refund cap.
  • The database is rebuilt before each run, so nothing persists into the next task.
  • The grader reads the database and the trail after the run, from outside the agent's process; the agent never sees the specification or the grader code.
  • The judge, when used, is a different model with the state difference in hand, so a persuasive reply has less to work with. Level 4 showed it can still be talked into an invented policy, which is why the specification is the next rung.
  • A code agent's equivalent is a test suite the agent never sees, run after it finishes. The quickest check of your own setup is one question: could the agent, with the tools it has, make the grader say pass without doing the task? If yes, it eventually will.

    Important: any grader the agent can reach, it will eventually satisfy without doing the job. Hacked-resolved rates of 28.57% fell to 0.56% when verification moved out of reach.
    What it means for us: the state the grader checks must be written by code the agent cannot edit, the grader must run after the agent and outside its process, and the tasks kept back for grading must stay out of the agent's view. Re-check all three every time the agent gets a new tool.
    Why it matters: optimising against an eval is what every tuning loop does, including the one that writes your prompt. The eval has to survive it.

    Carry forward from Level 7:

  • Reward hacking is measured at a quarter of "successes" when the grader is reachable and under 1% when it is not.
  • Policy in tools, a clean world per run, and a grader outside the agent's process are the three properties to keep.
  • A number from a benchmark the agent was tuned on is a number about that benchmark.
  • Level 8: making it run every day

    This level answers: how do you turn one eval run into a habit that catches regressions before customers do?

    The loop. Every change to prompt, tools, model, or harness runs the full task set, several trials each, and produces the same report: pass@k and pass^k by category with intervals, first-wrong-step counts, cost per pass, judge agreement. Nothing ships if a gate fails.

    Gates that make sense for an agent.

  • pass^k per category never drops below its floor, set a little under the last accepted run, by the width of that run's interval. For the example the last accepted values are 1.00 on refusals and 0.73 on actions.
  • Forbidden-tool attempts stay at zero on policy tasks.
  • Cost per passing episode stays under a ceiling per category.
  • Judge-versus-code agreement stays above its floor, or the judge's verdicts do not count.
  • The null agent still fails every action task, and every task is still solvable from the clean world. Re-run both whenever the task file changes.
  • Growing the task set from real traffic. Twenty hand-written tasks is a start. Anthropic's guide to agent evals recommends 20 to 50 tasks taken from real failures, and claims that CORE-Bench moved from 42% to 95% after grading corrections alone, which is the reason to fix the grader before growing the task count. The loop is: sample production traces, label the failures, turn each into a task with an expected state, re-run.

    Simulated customers. Some tasks need a conversation. The example re-ran 6 tasks that invite a follow-up question with a second model playing the customer from a hidden fact sheet. Over those tasks pass@3 was 0.83 and pass^3 0.67, and the agent asked a question in 0.11 turns per episode: it almost never asked, and resolved the ambiguity from the order list instead. Whether to ask is a behaviour to grade.

    Real sessions as the final grader. Once the agent is live, route a small share of real traffic to each new version, a canary, and grade it by state. SWE-chat (2026), 6,000 real coding-agent sessions, offers two signals to track there: whether the agent's change survived to the end of the session (44.3% did) and whether the user pushed back (in 44% of sessions). Both are measured on real work, not on a task set.

    What to log, per episode. Task id and trial. Prompt and tool versions. Every tool call with arguments, result, and wall time. The model's reasoning when it produced any. The final reply. The database before, after, and the difference. Tokens per call. The grade, check by check. The judge's verdict and reason. Anything less, and Level 3 cannot be done after the fact.

    Important: an agent eval is a habit with a report, not a number with a date.
    What it means for us: gate on pass^k per category, zero forbidden attempts, cost per pass, and judge agreement; grow tasks from real failures; re-run the null-agent and solvability checks whenever the task file changes.
    Why it matters: by Anthropic's account CORE-Bench moved 53 points from grading fixes alone. Check the grader before the agent.

    Carry forward from Level 8:

  • Gate on reliability per category, forbidden attempts, cost per pass, and judge agreement. Diff the report every run.
  • Grow tasks from real failures, 20 to 50 to start, and fix the grader before adding tasks.
  • Log the whole trail and the state difference, or the failure analysis cannot happen later.
  • The tools and approaches, observed

    This section is an observation, not a recommendation list. The graders and estimators were run on the example; the frameworks and benchmarks were read in their documentation and papers while preparing this article.

    Graders and estimators

  • State-and-path code grader (the example's). Best at: Exact, free, unreachable from inside the agent; grades end state, required and forbidden tools, step budget, reply keywords. What to watch: Blind on "change nothing" tasks without the path checks; only as good as the gold specification.
  • pass@k and pass^k (Sierra, τ-bench). Best at: Capability and reliability from the same n trials, exact for any k up to n. What to watch: 3 trials minimum, 5 to be smooth; intervals must resample tasks, not episodes.
  • Null-agent and solvability audit (the ABC checklist, UIUC and others). Best at: Catching graders that cannot fail and tasks that cannot pass, in minutes. What to watch: Re-run whenever tasks change; a null agent does not catch an unsolvable task.
  • First-wrong-step labelling (MAST, τ-bench, MP-Bench). Best at: Turning failures into a ranked list of fixes. What to watch: Experts agree on the step 16.2% of the time; keep the categories few.
  • Model-as-judge with state evidence (Agent-as-a-Judge, GAUGE, AgentRewardBench). Best at: Tone, reasoning quality, tasks without a written end state. What to watch: Over-estimates success, over-blames the model, favours its own family; report kappa against the code grader every run.
  • Frameworks

  • Inspect AI (UK AI Security Institute, open source). Best at: Agents in sandboxes, tool-calling solvers, a log viewer for trails. What to watch: Code-first; trajectory scoring is yours to write.
  • promptfoo (open source, Node.js). Best at: Declarative trajectory assertions out of the box: tool used, tool arguments match, tool sequence, step count, goal success. What to watch: Not Python; the assertions check the trail, not the state, unless you add a state check.
  • LangSmith (hosted and self-hosted). Best at: Production trace capture and the trace-to-task loop. What to watch: The grading is yours.
  • Public benchmarks, and how to read them

    A leaderboard number transfers to you along three axes: the grading method (final answer, final state, hidden tests), the repetition (one run or pass^k), and the distance from your domain. None gives the number you will get on your own tasks.

  • τ-bench and τ²-bench (Sierra). Best at: State-graded policy following with pass^k on 115 retail and 50 airline tasks, plus a second acting party in τ²; the closest public relative of a support agent. What to watch: A do-nothing agent passes 38% of airline tasks.
  • Terminal-Bench 2.0 (Stanford and Laude Institute). Best at: 89 container tasks graded by tests on the final state, 5 runs each, with intervals and cost per task. What to watch: One task can cost 100 million tokens.
  • SWE-bench Verified (OpenAI). Best at: Hidden-test grading on 500 real issues, screened by human annotators according to OpenAI. What to watch: The ABC audit found 5.2% of tasks pass with no fix and 24% of a top-50 sample wrong on inspection.
  • GAIA, WebArena, OSWorld (Meta, CMU, HKU). Best at: Showing the human-agent gap at launch, with humans above 70% and agents below 16% on all three. What to watch: String-match or environment-script grading; a Chrome version change moved OSWorld by 28 points.
  • BrowseComp (OpenAI). Best at: 1,266 hard verifiable web-research questions. What to watch: OpenAI's own report puts the agents' confidence far above their accuracy; the agent's confidence is not evidence.
  • The ABC checklist (UIUC and others). Best at: Auditing any benchmark, public or yours, for task validity and outcome validity; 7 of 10 audited benchmarks had grader flaws. What to watch: Applied to the example, it found 2 unsolvable tasks in 20.
  • A decision guide

  • Did the agent do the task?. Use: state checks plus required and forbidden tools plus step budget. Then check: the null agent fails it and a clean world can pass it.
  • Will it do it again?. Use: pass^k with k of 3 to 5, by category, task-clustered intervals. Then check: the paraphrased task set gives a number inside the same interval.
  • Why did this episode fail?. Use: first wrong step from the stored trail, one category per failure. Then check: the top category explains most failures; fix that one first.
  • Was the reply good?. Use: a judge given the trail, the state difference, and the specification. Then check: kappa against the code grader above 0.6, and no same-family judge.
  • Is it worth the money?. Use: tokens and seconds per passing episode by category. Then check: the maximum-to-minimum ratio across tasks, and the single-agent baseline at equal tokens.
  • Which agent in the pipeline lost it?. Use: stored briefs at every hand-off, graded end to end. Then check: hand-off loss, actor error, or grader error, from the trail, never from the agents.
  • Could the agent be gaming us?. Use: the three-property check: grader state out of reach, grader runs after and outside, graded tasks stay out of the agent's view. Then check: hacked-resolved rate on a sample read by a human.
  • Is it safe to let it act for real?. Use: zero forbidden attempts on policy tasks across all trials, and a canary on sampled traffic graded by state. Then check: the change-survival and pushback rates on real sessions.
  • Fifteen things to carry away

  • An agent can be wrong in its words, its actions, their order, their count, and what it left behind. Grade all five, the words last.
  • A reply-only grader passed 58 of 60 episodes and a do-nothing agent on every task. Run the do-nothing agent through every grader.
  • Check every task against the clean starting world. Two of the example's 20 tasks could not be passed by any agent, and no pass rate would have said so.
  • Report capability (pass@k) and reliability (pass^k) from the same runs, with intervals from resampling tasks. The gap is what customers experience.
  • Three trials is the floor, five is the recommendation.
  • Break every number out by task category. The example's 0.75 was a 1.00 on refusals and a 0.00 on escalations.
  • Vary the wording as well as the seed. Nearly half of equal paraphrases changed a code model's output.
  • Walk every failure back to its first wrong step and count the kinds. One kind, "replied before acting", explained every real failure in the example.
  • Count steps, tool errors, refused change attempts, and empty look-ups on every run. They move before the pass rate does.
  • A judge grades the evidence it is given. The example's transcript-only judge reached a kappa of −0.24 with the code grader; the state difference fixed its false alarms, not its misses.
  • Judges over-estimate success and over-blame the model; rules under-count passes. Report both, and never let the agent's own model judge it.
  • Price the successes, not the episodes. Cost per passing episode by category is the number; the field measured a 30-fold spread for the same task.
  • Every hand-off between agents is a lossy channel: about 8.5 points per prose hand-off, 2.8 per structured one. Keep a single-agent baseline at equal tokens.
  • Any grader the agent can reach, it will eventually satisfy without doing the job. Policy in tools, a clean world per run, a grader outside the process.
  • Make it a habit with a report: gates on pass^k per category, zero forbidden attempts, cost per pass, judge agreement; tasks grown from real failures.
  • References

    Agent benchmarks and audits

  • Mialon et al. (2023). GAIA: a benchmark for General AI Assistants. arXiv:2311.12983.
  • Zhou et al. (2023). WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854.
  • Xie et al. (2024). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv:2404.07972.
  • Yao et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045.
  • Barres et al. (2025). τ²-bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982.
  • OpenAI (2024). Introducing SWE-bench Verified. openai.com/index/introducing-swe-bench-verified.
  • Merrill et al. (2026). Terminal-Bench 2.0. arXiv:2601.11868.
  • Wei et al. (2025). BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents. arXiv:2504.12516.
  • Zhu et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825.
  • Sahoo et al. (2026). AgentLens: Revealing the Lucky Pass Problem in SWE-Agent Evaluation. arXiv:2605.12925.
  • Failure analysis and attribution

  • Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657.
  • In et al. (2026). Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation (MP-Bench). arXiv:2603.25001.
  • Raj et al. (2026). Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures. arXiv:2607.28802.
  • Rabanser et al. (2026). Reliability versus accuracy of language models over time. arXiv:2602.16666.
  • Mastropaolo et al. (2023). On the Robustness of Code Generation Techniques: An Empirical Study on GitHub Copilot. arXiv:2302.00438.
  • Judges

  • Zhuge et al. (2024). Agent-as-a-Judge: Evaluate Agents with Agents. arXiv:2410.10934.
  • Lù et al. (2025). AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories. arXiv:2504.08942.
  • Bodhwani, Tran, Wei (2026). GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents. arXiv:2609.12191.
  • Cost, hand-offs and multi-agent systems

  • Bai et al. (2026). How Do AI Agents Spend Your Money? arXiv:2604.22750.
  • Sayed Ali et al. (2026). The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI. arXiv:2607.06906.
  • Ao, Gao, Simchi-Levi (2026). On the Reliability Limits of LLM-Based Multi-Agent Planning. arXiv:2603.26993.
  • Chacon Sartori (2026). The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents. arXiv:2603.24284.
  • Tran, Kiela (2026). Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets. arXiv:2604.02460.
  • Kim et al. (2025). Towards a Science of Scaling Agent Systems. arXiv:2512.08296.
  • Reward hacking

  • Wang et al. (2026). The Verification Horizon: No Silver Bullet for Coding Agent Rewards. arXiv:2606.26300.
  • Zhang et al. (2026). Harnessing Agentic Evolution (AEVO). arXiv:2605.13821.
  • Lu et al. (2026). The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? arXiv:2606.04455.
  • Xia et al. (2026). RRSI: Regularized Recursive Self-Improvement of Agent Harnesses. arXiv:2609.24972.
  • Practice

  • Anthropic (2026). Demystifying Evals for AI Agents. anthropic.com/engineering.
  • LangChain (2026). LangSmith: evaluating agents. docs.langchain.com.
  • Baumann et al. (2026). SWE-chat: Coding Agent Interactions From Real Users in the Wild. arXiv:2604.20779.
  • Inspect AI documentation, inspect.aisi.org.uk. promptfoo documentation, promptfoo.dev/docs.
  • Actions
    What You Can Do
    • Export as PDF or Markdown
    • Batch Export to Notion
    • Bookmark & Highlight
    • Screenshot Tweet
    Create Free Account

    Includes 7-day Premium trial

    Advertisement