A harness for AI coding agents near production
Built project skills, defined subagents and guardrails so AI agents could diagnose production safely, with a human approving every fix.
- Role
- Sole designer and operator of the process
- Period
- August 2026 to present
- Status
- In use on a live production project (as of 29 September 2026)
- Outcome
- 11 project skills, 3 agent definitions and 168 tracked issues with root causes.
01 / Problem
The repository had a second author: an AI app builder that committed and deployed code without passing through my checkout. Early on, my local copy was behind production, which ran a function that did not exist on my disk. Agents also lost their context at every compaction, and a paid campaign was about to point real traffic at the system. I needed agents that could diagnose a live system in parallel and still could not break it.
02 / System
Select a component to read what it does and how it fails.
- Analysis agents. Each reproduces one issue, read-only against production, and writes one root-cause file. Parallel runs collided in a shared browser until each got its own context.
- Human approval. I read each root cause and approve a plan before code is written. The gate sits where blast radius changes, and I am the bottleneck.
- Fix agent. Plans first, then implements in its own git worktree after I approve. A stronger reviewer model is consulted before and after; I check its claims.
- Promotion chain. Fixes move through a chain of branches; the app builder's branch feeds staging. Agents cannot deploy infrastructure, and every task starts with a sync.
- Issue tracker. Anything undecided or broken lands here with root cause and plan, tagged found by AI or by human. Memory holds only per-machine facts.
03 / Decisions
Split diagnosis from fixing
- Decision
- Separate defined agents: analysis is read-only and runs in parallel, and fixing is plan-only until I approve, in its own worktree.
- Rejected
- Spawning ad hoc agents and tightening their prompts, after they started stepping on each other.
- Why
- Diagnosis is safe to parallelise and fixes are not. A live campaign raised the stakes, so the split sits where blast radius changes.
- Cost
- Every fix waits for my approval, and each agent definition needs maintaining.
Gate on a stronger model, then check it
- Decision
- I consulted a stronger reviewer model before and after implementing and before anything irreversible, and verified its claims against the data.
- Rejected
- Deferring to it, or skipping it to save quota.
- Why
- It caught a cross-tenant leak, a false claim in a vendor report and a hardening step that would have silently stopped load balancer logs. It also flagged an authorization risk in two functions that already enforced auth, so I did not act on it.
- Cost
- Usage limits. I hit the weekly cap mid-test-pass, so mechanical agents moved to a cheaper model and only real diagnosis escalated.
Keep knowledge in files, not chat
- Decision
- Rules, procedures and open problems live in versioned files. Each skill traces to a dated incident, and memory holds per-machine facts only.
- Rejected
- Leaning on session memory and long chats.
- Why
- Agents lose their working context at every compaction, and the files survive it. The sync-first skill exists because my copy fell behind production.
- Cost
- A ritual before every compaction: update skills and memory, then compact. And discipline about which file a fact belongs in.
04 / What broke
Agents racing in one browser
- Symptom
- Logging in as admin in one tab redirected another tab's signup page into the admin app.
- Cause
- Every tab shared one browser context, so one cookie jar.
- Fix
- Named, isolated contexts held in a registry on the automation server. I confirmed it persists across subagent calls before relying on it.
One agent's cleanup nearly erased another's work
- Symptom
- A fix agent's post-commit cleanup ran a checkout that reverted another agent's uncommitted changes. They survived only because that agent kept writing.
- Cause
- Two agents shared one checkout.
- Fix
- Each agent now commits from its own worktree and never discards changes it did not create.
A debugging loop
- Symptom
- The agent kept trying fixes for the auth crash loop and none held.
- Cause
- It was fixing before diagnosing.
- Fix
- I stopped it and reset the method: use the official docs actually in use, get the reviewer's read, read the logs as far as they go, log every attempted fix, then diagnose. The root cause followed.
05 / Outcome
- Nearly every skill and agent definition was written after a specific incident and is dated to it.
- Across two sessions from 16 August to 29 September 2026, I directed 55 of 155 reviewer-model consultations, and 20 of 55 subagent spawns came from my own definitions.
- These figures describe practice, not results. I have no speed, quality or cost measurement, so I quote none.
06 / The rule I took from this
Stack
Next
Working on something similar? Email me about this build.
Related: An AI voice front desk for dental clinics, A production AWS backend on HIPAA-eligible services.
Hiring for this kind of work? AI backend engineer.
Next case study: An operations console for a healthcare integration engine.