Applied AI · 2026 to present

Skills and model routing for AI coding agents

Skills that hand agents current third-party docs, builds, tests and browser sessions, with each kind of task sent to the right model.

The problem

What was needed, and what made it hard.

An AI coding agent writes code from what it remembers, and for a third-party API that memory can be a year old. Searching the web on every task to make up for it costs time and session quota, and so does running the strongest model for routine commands. I wanted agents that work from current docs, run the project's builds and tests the same way every time, and spend the most capable model only where a mistake costs the most.

The system

How one task moves from plan to a reviewed build, and which model does each step.

Initial documentation 1 Plan per task 2 Plan review 3 Implementation 4 Routine runs 5 Final overview 6 Initial documentation 1 Plan per task 2 Plan review 3 Implementation 4 Routine runs 5 Final overview 6
How one task moves from plan to a reviewed build, and which model does each step.

Select a component to read what it does and how it fails.

  1. Initial documentation. The most capable model (Opus) writes the first project documentation, so every later task starts from the same written ground.
  2. Plan per task. The mid-tier model (Sonnet) plans each task on its own, reading the project skills and the downloaded API references instead of searching again.
  3. Plan review. The most capable model reviews that plan before any code is written. This is where a wrong assumption is cheapest to catch.
  4. Implementation. The mid-tier model writes the code. Small inline fixes go to the smallest model (Haiku).
  5. Routine runs. The smallest model handles pull, push, builds, reading logs, user-acceptance checks and research, through skills that run each one the same way.
  6. Final overview. Once the build compiles without errors and the test suites pass, the most capable model reviews the whole implementation.

Decisions

What I chose, what I turned down, and what it cost.

1Match the model to the task

I chose
The most capable model for initial documentation, reviewing each plan and the final overview; the mid-tier model for planning and coding; the smallest for inline fixes and routine runs.
I turned down
One model for everything, either the strongest (quota runs out) or the cheapest (mistakes reach review).
Why
Session quota is finite. Pull, push, builds, logs and user-acceptance checks do not need deep reasoning, while a plan or a finished implementation is where an error is most expensive.
What it cost
More handoffs, and a cheaper model can miss something, so the build, the test suites and the final review have to catch it.

2Fetch the docs once, then read them locally

I chose
A skill per third-party API that explains how to call it, with its documentation fetched once from the live reference and kept as the skill's resources.
I turned down
A web search on every task, or trusting what the model remembers about the API.
Why
One search and a download replace repeated searches, and the agent reads current docs instead of guessing from a knowledge cutoff.
What it cost
The local copy can fall behind the vendor, so a skill has to fetch the live reference again when the API changes.

3Turn routine commands into skills

I chose
Skills for running local builds, running the test suites and keeping a persistent Playwright session, so new work is checked against what already works.
I turned down
Letting each agent rediscover the commands and open a fresh browser session through the Playwright MCP tools for every check.
Why
The same steps run the same way every time, the test suites guard existing behaviour, and a persistent session needs fewer Playwright MCP tool calls.
What it cost
The skills have to be kept in step with the project as its build and test setup change.

What broke

What I noticed, what caused it, and how I fixed it.

A payment API the model thought was new

What I saw
An AI agent built a Stripe payment integration on an API version that was later deprecated, and the model never flagged it.
The cause
Its knowledge of that API was about a year old, so it took the version as recently introduced and unlikely to be deprecated soon.
The fix
I wrote a Stripe skill that fetches the live reference first. Working from it, the agent caught the deprecation and the integration was fixed.

Outcome

What came of it, and what it was built with.

  • Work against a third-party API starts from a skill and its downloaded reference: one search, then local reads.
  • The Stripe skill's live reference caught a deprecated API version that the model's own knowledge missed.
  • Builds, test suites and browser checks run through skills, and agents reuse one persistent Playwright session.
  • I have not measured accuracy, cost or quota savings, so I quote no figures. These describe the practice.

Built with

  • Claude Code
  • OpenCode
  • Project skills
  • Claude Opus
  • Claude Sonnet
  • Claude Haiku
  • Playwright
  • MCP servers
  • Test suites
  • Stripe API

The rule I took from this

The line I keep from this project.

Give the agent today's docs, not its memory.

Rule 10 of 10

Keep reading