AI Agent Evaluation That Reads the Transcript, Not the Summary
An agent’s final message is the least reliable part of its trace. "Done — I’ve refunded the order and emailed the customer" is what you get whether the refund succeeded or returned a 409 three times in a row. Grading the summary measures the agent’s confidence, not its work.
The trace itself — tool calls, arguments, results, status codes — is the evidence. It is also structured text, which is exactly what a decision model is for: did this reach the goal, what broke, how many steps were wasted, did it tell the user something untrue.
- Grade every run, not a sample
- Failure taxonomy as declared labels
- Confidence on every answer
- Free output tokens
The problem
The Job: Find the Runs That Failed While Claiming Success
Agent evaluation splits into two questions. Outcome: did the world end up in the state the task asked for? Process: did it get there without a retry loop, a hallucinated parameter or a step that will have to be undone? Only the first can be checked with an assertion, and only sometimes — most production tasks do not have a clean end-state to diff.
The process questions are what you actually debug against, and they are judgement calls over a transcript. "Repeated a failing call without changing anything" is obvious to a reader and awkward to express as a rule, because whether a retry is reasonable depends on the error it got back.
The specific failure worth spending money to catch is the false completion: the run that reports success for something that never happened. It does not show up in error rates or latency dashboards, it shows up in a support ticket. Grading the trace catches it; grading the summary is how it got past you in the first place.
Why a decision model
Why a Decision Model Beats Asking an LLM to Review the Trace
Agent traces are long and there are a lot of them. Both facts push against a grader that writes prose.
A failure taxonomy you declare is a failure taxonomy you can count
Wrong tool, ignored precondition, retry loop, hallucinated parameter, false completion. Declared as labels, they aggregate into a chart that says which failure mode grew this week. Written as free text, they aggregate into nothing.
Traces are long, and you pay for input only
A transcript is the expensive part of the prompt, and a prose reviewer adds its own long output on top of it. Here the output is free, so the cost of grading a run is the cost of reading it once.
Cheap enough to grade every run in CI and in production
At the price measured below, grading a thousand runs costs a few cents. That turns agent evaluation from a release ritual into a continuous signal you can alert on.
Confidence tells you which traces a human should read
The runs the grader is unsure about are almost always the interesting ones — novel failure modes it has no good label for. Sorting by confidence is a far better review queue than sorting by timestamp.
The outcome and the process questions share one read
Completion, failure mode, efficiency and user-visible harm are four questions about the same transcript. Asked together they cost one pass over the trace.
Grade a Run That Claims It Succeeded
A refund agent that hit the same 409 twice, invented a `force` parameter, emailed the customer anyway and signed off with "Done". Run it and see whether the grader agrees with the agent.
Did the run reach the goal, and where did it go wrong?
1Text
Agent transcript1,106 / 100,000
2Questions
4 in this requestEditing is free. Sign in and you come straight back here with your text and questions — no card needed.
3Answers
Answers appear here
Each answer returns a probability for every option and a confidence score.
In code
What You Would Write Next
The grader only sees what you serialise. Include the tool results — a trace of calls without their responses cannot distinguish a success from a 409.
// Serialise the trace the way a reader would see it: the task, then the
// tool calls with their results, then the final message. The grader needs
// the results - a trace of calls without responses cannot show a false
// completion.
function renderTrace(run: AgentRun) {
return [
`TASK GIVEN TO THE AGENT\n${run.task}`,
"TRANSCRIPT",
...run.steps.map(
(s, i) => `${i + 1} ${s.tool}(${s.args}) -> ${s.status} ${s.result}`
),
`FINAL: ${run.finalMessage}`,
].join("\n");
}
const { answers } = await jev.decide({
state: renderTrace(run),
questions: {
task_completed: {
type: "noul",
instructions:
"Was the task actually completed? Judge the effect of the tool calls, not what the final message claims.",
},
failure_step: {
type: "choice",
instructions: "Select what went wrong in this run.",
criteria: {
none: "Did what was asked, no wasted or harmful steps.",
retry_loop: "Repeated a failing call without changing anything.",
hallucinated_parameter: "Invented an argument the tool does not have.",
false_completion: "Reported success for an action that never succeeded.",
ignored_precondition: "Read a constraint and then acted against it.",
},
},
},
});
// The alert that actually matters: the agent said done, the trace says no.
if (answers.task_completed.noul < 0.5 && run.finalMessage.match(/done|completed/i)) {
alerts.falseCompletion(run.id, answers.failure_step.probabilities);
}The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.
Real cost
One Run Costs $0.000038
Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the agent transcript in the playground above, with all four questions in one call.
It checks out against the list rate of $0.042 per million input tokens: 901 ÷ 1,000,000 × 0.042 = $0.000038. The 135 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.
- Input tokens
- 901
- One run
- $0.000038
- 1,000 runs
- $0.038
- 1,000,000 runs
- $37.84
The only thing charged
Round trip 0.50s from a laptop, network included
Same questions, same length of input
At the model list rate, before our margin
Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.
Questions
Agent Evaluation: The Questions People Actually Ask
Does this replace assertions and integration tests?
No. Where you can assert on a real end state, assert — it is cheaper and exact. This covers the majority of agent behaviour that has no assertable end state: whether the path was sane, whether the summary was honest, whether a step will have to be undone.
How long a trace can I grade in one call?
Up to the 32,000-token context. Longer runs are better graded in segments anyway — a per-segment verdict tells you where it went wrong, where a single verdict over fifty steps tells you only that it did.
What should I put in the state?
The task, the tool calls with their arguments, the tool results including error bodies and status codes, and the final message. The results are the part people leave out and the part that makes false completions detectable.
Can I grade runs in CI on every pull request?
Yes, and the cost is why. A hundred-run suite costs well under a cent at the figure measured on this page, so the grading step is not what makes your pipeline slow or expensive — the agent runs themselves are.
Will it catch a failure mode I did not put in the label list?
Not as a label — a choice question can only return labels you declared. It shows up as low confidence and a spread distribution instead, which is why sorting by confidence is the right way to find the failure modes your taxonomy is missing.
Does a grading call ever change my agent’s behaviour?
No. It is a separate read-only call over a transcript you already have. Grade asynchronously off your trace store if you do not want it on the request path at all.
Related
Other Decisions in This Shape
LLM as a judge
Grade generated answers against a rubric and get sortable scores instead of a paragraph of praise.
RoutingLLM router
Decide which model, tool or workflow should handle a request, with a probability you can threshold.
RetrievalRAG evaluation
Check retrieval quality and answer grounding chunk by chunk, cheaply enough to run it on everything.
Run It on Your Own Data
Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.
