Guide · 6 chapters
Building production AI agents
Tool use against real APIs, eval harnesses, observability, and the kill-switches you'll want.
Chapter 01
What 'production' means for an agent
A production agent is one your team is comfortable letting run unattended against real customer data. That bar is much higher than 'demo works.' This guide is about what 'comfortable' means, mechanically.
Chapter 02
Tools, typed and bounded
Every tool the agent can call is a typed function with an explicit schema, narrow inputs, and a per-call budget. Tools are versioned. New tool versions ship behind a flag and get evaluated before they go live for the agent.
Chapter 03
The eval suite
Versioned set of graded transcripts that runs on every prompt change. We aim for 200-1500 transcripts per agent, refreshed weekly from production. Pass rate floor is 92%; PRs that drop any axis below that are blocked.
- Versioned in the same repo as the agent.
- Refreshed from real production transcripts weekly.
- Run on every PR that touches prompt, tools, or model.
- Failure modes graded across correctness, tone, refusal, cost.
Chapter 04
Observability — prompt, tool, latency, cost
Every agent conversation is fully traced: the prompt sent, the tools called, the responses received, the tokens used, the cost computed. Traces are sampled into a queryable store so engineering can grep for failure modes without re-running the agent.
Chapter 05
Kill-switches and budget caps
Every agent has a hard kill-switch wired into the loop, not the dashboard. Per-conversation budget caps prevent a runaway loop from costing $50. Per-tool budgets prevent a single tool from being called 100 times in one conversation.
Chapter 06
Human handoff
When the agent escalates (policy says so, eval pattern says so, customer asks), the human gets the full conversation context, not a summary. Handoff is reversible — humans can hand back to the agent if the customer is happy.
More guides
Multi-tenant SaaS, end to end
Postgres RLS, RBAC, metered billing, and a SOC 2-ready audit trail — wired in before the first tenant signs.
Core Web Vitals — the playbook we run
LCP, INP, CLS — how we diagnose, fix, and lock in. With the GitHub Action we use to enforce.
WordPress to Next.js without losing SEO
Redirect maps, content migration, hreflang, and the post-launch monitoring that catches misses.
Run this with your team
Book the workshop version.
A half-day workshop with your team — same content, your codebase. We tailor the chapters to where your team is today.