Research
Notes from
the lab.
What we learn building voice and agent systems for real operations, written by the people who build them.
Full archive · 12 notesWhat AI Agents Can Reliably Own in Production
The useful question is not whether an agent can complete a demo. It is which bounded parts of real work it can own repeatedly, with tools, approvals and recovery in place.
Read the note →Recent notes
18 Sept 2026InfrastructureMCP Is Not the Product: What Secure Tool Connections Actually ChangeMCP gives agents a standard way to reach tools and data. The business value still comes from choosing the right work, limiting permissions and designing the approval path.4 min→16 Sept 2026EvaluationHow to Evaluate an AI Agent Before It Touches Real WorkA production evaluation should test outcomes, tool use, refusals, handoffs and recovery. A polished demo and a list of model benchmarks are not enough.4 min→24 Jun 2026Voice systemsAI Voice Agents for Home Services: Never Miss a JobAI voice agents for home services answer every call 24/7 — even when your techs are on a roof or under a sink — triage the emergency, and book the job fast.8 min→22 Jun 2026Voice systemsAI Voice Agents for Dental Clinics: Book Patients 24/7AI voice agents for dental clinics answer every call 24/7, qualify new patients, and book into your practice software — so the front desk stops losing patients to voicemail.8 min→19 Jun 2026Voice systemsAI Voice Agent vs IVR vs Voicemail: An Honest ComparisonAI voice agent vs IVR vs voicemail: an IVR routes calls and voicemail stores them, but only an AI voice agent answers, qualifies, and books. Here's the honest pick.8 min→17 Jun 2026Voice systemsAI Voice Agent vs Human Receptionist: Which Do You Need?AI voice agent vs human receptionist: a human wins rare complex calls, but AI answers every call 24/7 and rings leads back in seconds. Here's the honest pick.7 min→
Subjects: voice systems, agent operations, evaluation, governance, infrastructureBrowse the archive →
Apply the research
Bring one process to a working session.
Thirty minutes with the people who build the systems. We map the work and tell you honestly whether an agent should do it.