An AI voice front desk for dental clinics
Took an AI receptionist for dental clinics from AI-scaffolded code to production, with three voice providers and per-practice tenant isolation.
- Role
- Owner of the product's engineering
- Period
- August 2026 to present
- Status
- Live in production (as of 29 September 2026)
- Outcome
- Launched to production on 3 September 2026; the AWS backend I designed took over in mid-September.
01 / Problem
A dental practice wants an AI receptionist that books, reschedules and cancels appointments and checks insurance eligibility. It has to sit on top of the practice's existing scheduling system, serve many practices from one deployment, and keep each practice's data strictly apart. I joined a codebase scaffolded by an AI app builder that also committed and deployed code, so production could run things that never touched my checkout. My job was to make it safe to launch, then keep it running.
My role: Owner of the product's engineering: backend, infrastructure and the internal admin console
02 / System
Select a component to read what it does and how it fails.
- Voice provider. Answers the call. One switch picks among three providers, each with verified webhooks. Each is its own pipeline, so fixes must land in all three.
- Tool endpoints. The agent's only actions: booking, eligibility, lookup. It once re-derived a time from its own words, dropping the offset: 12:00 PM became 5:00 AM.
- Application backend. Holds the business rules and per-tenant scoping. A query missing its tenant filter leaks across tenants, as one total once showed.
- Tenant database. Postgres with row-level security per practice. Missing grants show up as permission errors; a missing filter shows up as nothing, which is worse.
- Sync jobs. Scheduled jobs sync appointments, providers and time zones with the practice's scheduling system. A per-practice circuit breaker pauses a practice when its connector goes down.
03 / Decisions
Keep the product narrow
- Decision
- An AI layer on the practice's own scheduling system, which stays the system of record.
- Rejected
- Building records, billing or claims features into the product.
- Why
- The practice already runs on that system, so the product only has to be right about calls, bookings and eligibility.
- Cost
- It depends on that system's API and on each connector staying up. When one practice's connector went down, a queue ran away.
Make provider identity per practice
- Decision
- Each practice holds its own credentials for the eligibility API, and I removed the environment-variable fallback.
- Rejected
- One shared credential as a default for every practice.
- Why
- With a fallback, one practice's credentials could serve another. This fixes isolation at the provider boundary, not only in the data layer.
- Cost
- Every practice needs its own setup before eligibility checks work.
Self-host auth instead of waiting
- Decision
- I stood up a self-hosted auth stack on the staging server, and 10 of 10 raw requests then passed.
- Rejected
- Waiting for the vendor to fix a known upstream bug.
- Why
- I sent 25 raw requests straight at the managed service, bypassing my own code. All 25 failed with a "token issued in the future" error while the server's clock header advanced normally.
- Cost
- I now operate an auth service that a vendor used to run for me.
04 / What broke
A total that did not add up
- Symptom
- An insurance carrier count on a screen I was redesigning read 80, and a reviewer model doubted it.
- Cause
- It was the exact sum of several tenants' rows. The query had no tenant filter, at a call site an earlier leak fix had missed.
- Fix
- I added the filter and checked it against the isolation tests.
A queue that would not stop growing
- Symptom
- After one practice's connector went down, failed-job rows passed 4,182 and were still climbing.
- Cause
- The scheduler advanced its last-pull time only on success, and the error was not retryable, so the already-queued guard never fired.
- Fix
- I tracked the last attempt as a backstop and added a per-practice circuit breaker. Production held flat at 4,519 for over 2.5 hours. I had also claimed a status still read "connected" without checking production. I checked, it was wrong, and I corrected the record.
A date the agent never questioned
- Symptom
- Asked for next Tuesday, the voice agent searched a window years in the past. The slot tool correctly found nothing, and the agent never questioned it.
- Cause
- The rendered prompt never stated today's date.
- Fix
- I reproduced it live, then fixed all three providers and both prompt paths, live-rendered and baked in.
05 / Outcome
- It launched to production on 3 September 2026, and the AWS backend I designed became its production backend in mid-September.
- An early audit found no gap behind an alarming lint count, and showed the behavioural isolation tests covered only 6 of about 90 tenant-scoped tables.
06 / The rule I took from this
Stack
Next
Working on something similar? Email me about this build.
Related: A production AWS backend on HIPAA-eligible services, A harness for AI coding agents near production.
Hiring for this kind of work? Python backend engineer.
Next case study: A social and market signal intelligence backend.