The problem
Payment support mixes simple policy questions with private account issues. I wanted the agent to answer the first kind quickly and check who it was talking to before handling the second.
What I built
Voice and chat use the same Claude Agent SDK agent. I gave it seven MCP support tools and added six checks before a response reaches the user. Vapi handles voice, and Supabase stores the conversations, tool calls, tickets, and evaluations.
When a person needs to step in
The handoff books a Cal.com callback during support hours, emails the support inbox, and posts to Discord. If the calendar call fails, the system records the failure. The agent can’t tell someone they have a booking when it didn’t make one.
What I tested
The 2 October records show a 4.2-second median voice turn on Render and eight of eight brief scenarios passing. Those are results from that run, not a promise that every future call will be that fast.
What testing caught
My first retrieval tests passed with fixture vectors. Real searches showed what those tests missed. I separated the test data from the live knowledge base so that a passing test actually means something.
Recorded evaluation · 2 October 2026. The figures here come from those test records.