Skip to content
diviteb

Guide · 6 chapters

Building production AI agents

Tool use against real APIs, eval harnesses, observability, and the kill-switches you'll want.

Chapter 01

What 'production' means for an agent

A production agent is one your team is comfortable letting run unattended against real customer data. That bar is much higher than 'demo works.' This guide is about what 'comfortable' means, mechanically.

Chapter 02

Tools, typed and bounded

Every tool the agent can call is a typed function with an explicit schema, narrow inputs, and a per-call budget. Tools are versioned. New tool versions ship behind a flag and get evaluated before they go live for the agent.

Chapter 03

The eval suite

Versioned set of graded transcripts that runs on every prompt change. We aim for 200-1500 transcripts per agent, refreshed weekly from production. Pass rate floor is 92%; PRs that drop any axis below that are blocked.

  • Versioned in the same repo as the agent.
  • Refreshed from real production transcripts weekly.
  • Run on every PR that touches prompt, tools, or model.
  • Failure modes graded across correctness, tone, refusal, cost.

Chapter 04

Observability — prompt, tool, latency, cost

Every agent conversation is fully traced: the prompt sent, the tools called, the responses received, the tokens used, the cost computed. Traces are sampled into a queryable store so engineering can grep for failure modes without re-running the agent.

Chapter 05

Kill-switches and budget caps

Every agent has a hard kill-switch wired into the loop, not the dashboard. Per-conversation budget caps prevent a runaway loop from costing $50. Per-tool budgets prevent a single tool from being called 100 times in one conversation.

Chapter 06

Human handoff

When the agent escalates (policy says so, eval pattern says so, customer asks), the human gets the full conversation context, not a summary. Handoff is reversible — humans can hand back to the agent if the customer is happy.

Run this with your team

Book the workshop version.

A half-day workshop with your team — same content, your codebase. We tailor the chapters to where your team is today.