A production AWS backend on HIPAA-eligible services
Designed and deployed a production AWS backend for a healthcare SaaS, checked service by service against AWS's HIPAA-eligible list.
- Role
- Designed and deployed it independently as an infrastructure-as-code app; deploys stayed human-run
- Period
- 2026
- Status
- Live in production (as of 29 September 2026)
- Outcome
- Became the production backend in mid-September 2026.
01 / Problem
A healthcare SaaS ran on managed services: an edge runtime, plus a managed database and auth service. It needed a hosting path made only of services AWS lists as HIPAA-eligible, without losing tenant isolation, and with an auth stack the team controlled. A managed-auth outage had already shown what waiting on someone else's fix costs. My feasibility audit found the lock-in lopsided: leaving the edge runtime was an estimated 8 to 15 engineer-days, leaving the managed database and auth 60 to 100.
02 / System
Select a component to read what it does and how it fails.
- Regional WAF. Filters traffic in front of the load balancer. It stands where a CDN would have, and it is regional, so it guards only this balancer.
- Load balancer. Holds the certificate and listeners in front of the gateway. It was deleted once by a rollback that a bad container health check triggered.
- Gateway. Envoy routes to the auth, API and storage services and answers the ready endpoint that health checks now target. HTTP/1.1 passes, 1.0 gets a 426.
- Platform services. Self-hosted auth, REST and storage services on Fargate, for team-controlled auth. Only the storage client validates the database certificate chain, which broke replay until fixed.
- Database. RDS Postgres, single-AZ to hold cost. Migrations replay as one-off tasks; 18 of 185 need tables the auth and storage services create at boot.
03 / Decisions
Check the vendor's own list
- Decision
- I cross-checked every service against AWS's HIPAA-eligible reference, one by one.
- Rejected
- Building from memory of which services qualify.
- Why
- The list caught naming mismatches: the load balancer is listed as Elastic Load Balancing, and Security Hub appears only in its posture-management form. It also showed that third-party providers need their own business associate agreements.
- Cost
- Time before any code, and a list of obligations no code can close.
Move the edge to a regional WAF
- Decision
- When an account-level restriction blocked creating the CDN, I put a regional WAF on the load balancer.
- Rejected
- Stalling the edge layer on a support ticket.
- Why
- The WAF kept the protection I needed, and I designed the fallback so a CDN can be added later with nothing to tear out.
- Cost
- No CDN caching, which was already off in every planned behaviour, and no locking of the balancer behind a CDN. It was already internet-facing, so this is a deferral, not a regression.
Cut cost, and name each cut
- Decision
- A single-AZ database, one NAT gateway with a gateway endpoint, GitHub Actions for deploys and minimal WAF rules. Security Hub, Inspector and AWS Backup are deferred.
- Rejected
- Interface endpoints, CodePipeline and the fuller control set at pilot scale.
- Why
- At this scale interface endpoints were net-negative. My estimate is about 180 to 230 dollars a month for a pilot, against about 350 to 500 with the fuller controls.
- Cost
- No database failover, and deferred controls that remain open work.
04 / What broke
An auth crash loop on every boot
- Symptom
- The auth service crash-looped silently on every boot.
- Cause
- The connecting database role's search_path was never set. It governs the migration-tracking lookup and every unqualified query.
- Fix
- I set it live and wrote it into the bootstrap runbook.
A health check that deleted its own balancer
- Symptom
- Every task failed its container health check. The platform's own circuit breaker rolled back the update and deleted the new load balancer, certificate and listeners.
- Cause
- The check ran a tool that is not in the image. I found it by comparing what the container did with what the check tested.
- Fix
- I rewrote the check to use the gateway's ready endpoint.
A build hash that moved by itself
- Symptom
- The container image digest changed between two synth runs with no source change, which would have forced needless redeploys.
- Cause
- Session-local directories leaked into the build context. My first fix was incomplete, and excluding the whole infra directory broke the build.
- Fix
- I excluded subdirectories one by one and confirmed the digest stayed stable across an unrelated edit.
05 / Outcome
- The public domain was repointed to this deployment on 15 September 2026, and it was confirmed as the production backend the next day.
- All 185 migrations replayed end to end on the new stack.
- Eligible services are a starting point, not a finished control set: several controls were deferred and third-party agreements were still open.
06 / The rule I took from this
Stack
Next
Working on something similar? Email me about this build.
Related: An AI voice front desk for dental clinics, A harness for AI coding agents near production.
Hiring for this kind of work? Python backend engineer.
Next case study: A harness for AI coding agents near production.