Skip to content
diviteb
Illustrative

Example engagement showing how we approach this kind of project. The client, quotes, and figures are illustrative, not a real client record.

Customer support · AI agent

An agent that took 38% of ticket volume in week three.

Their support team was buried under 12,000 tickets a month. We deployed an agent wired to their order, refund, and CRM systems — with evals against six months of real transcripts and a kill-switch wired into the loop.

Outcome · week 3

shipped

38%

Of support tickets deflected by week three — zero false-positive escalations

AI SDKTool useEvalsTwilioLinear

38%

Ticket deflection in week three

$0.04

Median cost per resolved chat

0

False positive escalations to date

The starting state

The support queue was growing faster than headcount. Hiring a fourth tier of agents was on the table — and so was sunsetting the support SLA. We pitched a third option: an agent that handles the top 12 ticket categories, hands off cleanly when policy says so, and gets evaluated against real transcripts every release.

Ticket deflection · 8-week ramp

54% by W8

live
W1
W2
W3
W4
W5
W6
W7
W8
38% deflection by week three · 0 false positives

What we shipped

We treated the agent like a regulated system, not a chatbot. The eval suite ran on every prompt change against six months of real customer transcripts. Tools were narrowly typed and bounded by per-call budgets the agent could not escape.

  • 12 tools wired to the real APIs — refundOrder, lookupShipment, updateAddress — with signed schemas.
  • Eval suite of 1,400 graded transcripts; 92% pass rate required before any prompt change shipped.
  • Per-conversation budget cap of $0.50, hard kill-switch wired into the loop, full trace observability.
  • Human handoff with full conversation context piped into Linear — agents see what the bot saw.

Tool registry · last 30 days

Tool
Calls
OK
Fail
refundOrder
412
412
0
lookupShipment
1,840
1838
2
updateAddress
218
217
1
escalateToHuman
92
92
0
All tools signed · per-call budget caps live

Where it landed

Week three: 38% of incoming tickets handled end-to-end by the agent, with zero escalations marked as false positives. Median cost per resolved chat: $0.04, all tokens in. The hiring plan was scrapped; the support team now owns escalations, complex cases, and the eval suite.

Their agent took 38% of our ticket volume in week three and didn't escalate a single false positive. The eval suite alone was worth the engagement.
IllustrativeHead of SupportD2C subscription retailer

Ready when you are

Got a support queue you can't hire out of?

A 30-minute discovery call. We'll tell you which ticket categories an agent can take — and which need a human.