Skip to content
diviteb

AI · 2026-04-22 · 8 min

How we run an eval suite for production agents.

Demos pass on friendly inputs. Production breaks on adversarial ones. Here's the eval harness we wire into every agent — and what we look for in the failure modes.

ET

Engineering team

AI practice

The demo problem

An agent demo on Friday is a beautiful thing. The customer asks a question, the agent calls the right tool, the answer is correct. Everyone agrees: ship it. The agent goes to prod Monday and by Wednesday someone in support has flagged that it refunded $9,000 to the wrong customer.

The failure mode isn't that the agent is dumb. The failure mode is that the demo set was friendly. Production is adversarial — bored kids, edge-case business logic, customers who type in dialect, and the long tail of inputs no PM imagined.

What an eval suite is

An eval suite is a versioned set of graded transcripts that runs every time the prompt, model, or tool registry changes. For each transcript, you have an input (or full conversation), an expected outcome, and a grading criterion. The agent runs the input, you compare its output to the expected outcome with the criterion, and you get a pass/fail or a graded score.

The critical word is 'versioned' — your eval suite is part of the codebase. It's reviewed in PRs. New transcripts come from real production conversations (with PII scrubbed) every week.

  • Versioned in the same repo as the agent — eval changes get the same review as code changes.
  • Sized for signal, not breadth — 200 well-graded transcripts beat 2,000 vibe-checked ones.
  • Refreshed from real production transcripts weekly — the eval set learns what production sees.
  • Run on every PR that touches the prompt, the tool registry, or the model — no exceptions.

What we grade for

Pass/fail is too coarse for most agent behaviors. We grade across four axes: correctness (did it call the right tool with the right arguments), tone (did it match the brand voice we agreed on), refusal (did it correctly refuse when policy says so), and cost (did the conversation stay under the per-conversation budget).

Set a pass-rate floor on every axis before launch. Our default is 92%. If a prompt change drops any axis below the floor, the PR is blocked.

What it costs to run

Price it in tokens. Count the transcripts. Multiply by the average tokens per conversation, add the grader's tokens, and apply the model's list price. Do that arithmetic once before you argue about running it on every PR.

Then compare that bill to one bad refund, one leaked record, or one regression that reaches production. The eval suite is the cheapest line item in the agent's budget.

Run this in your team

Talk through it on a call.

A 30-minute discovery call. We'll walk you through how this would apply to your stack.