Dillon Green
ALL WRITING

WORKFLOW Applied AI in practice

How I ship production software by directing AI agents

Most of the code I ship, I didn't type. I'm still accountable for every line. Here's the workflow that makes both of those true.

The biggest thing I've built this way is an internal parts and maintenance platform for an airline simulator department. It tracks inventory, purchasing, repair work and budgets for a fleet of full flight simulators. It's about 72,000 lines, it's rolled out to a forty-person department and in daily use, and I'm the only developer. AI coding agents wrote most of it. I decided what it should be, how it fits together, what "done" means for each piece, and whether each piece ships.

None of what follows is exotic. It's mostly the habits of a careful tech lead, applied to a team that types very fast, never gets tired, and sometimes says untrue things with complete confidence.

Several agents, one tech lead

On a normal build day I have three or four Claude Code sessions running at once, each on its own task. One writes an importer for a legacy data export. One adds regression tests around a module I'm about to change. One fixes a layout bug on the phone view. The tasks are chosen so they don't overlap, and every agent's tests run against their own throwaway copy of the database, so they don't trip over each other.

My job is the part that doesn't parallelize: the data model, the boundaries between modules, the order things land in, and the final read of every diff. I don't write most of the code, but nothing merges that I haven't understood.

How the work gets split matters more than how many agents run. Tasks that touch the same tables or templates go to one agent, in sequence. Independent tasks run side by side. More agents on tangled work just produces merge conflicts faster.

Specs are goals, done-criteria and hard rules

I don't hand an agent "add a receiving screen." I give it a goal, a list of what done looks like, and the rules it can't break. A typical spec is short:

GOAL
  Receiving a purchase order puts the parts into inventory,
  at the bin the receiver picks, in one step.

DONE WHEN
  - A partial receipt leaves the order line open for the remainder
  - Every received unit writes one audit-ledger row
  - A new regression suite covers full, partial and over-receipt
  - The full suite passes; lint and invariant counts did not go up
  - Screenshots at 1440 and 390 px, light and dark

HARD RULES
  - Read data only from the scrubbed copy; never print real records
  - No schema change without asking first
  - Do not deploy, restart the service or touch backups
  - If the spec and the code disagree, stop and report

The done-criteria are what make review possible. "Done" isn't the agent saying it's done; it's a list I can check, one line at a time.

The hard rules draw a line between preparing and acting. Agents prepare. A human publishes, deploys and submits. An agent can write the migration, the release notes and the deploy script, but it doesn't run them against anything real. That rule has never cost me meaningful time, and it caps the worst case at a wasted afternoon.

Tests are the contract

What makes this work more than anything else is a test suite the agents can't argue with. On the platform that's more than 340 automated checks:

  • 70 regression suites that run in parallel, each against its own private copy of the database, so they can't interfere with each other or with anything real.
  • 138 end-to-end checks in headless Chrome, driven by a small Chrome DevTools Protocol client I wrote by hand. They log in, click through real workflows and check what ends up on screen.
  • A 30-thread concurrency test that hits the app and its SQLite database at once, to keep the choice of a single-file database honest.

On top of that sits a quality gate that only ratchets. It counts lint warnings and invariant violations, things like an orphaned row or a part whose quantity doesn't match its ledger, and remembers the numbers. A change can bring them down. It can't push them up.

# quality gate, simplified
baseline = load("gate_baseline.json")
current = {"lint": count_lint(), "invariants": count_invariant_violations()}

for key, now in current.items():
    if now > baseline[key]:
        fail(f"{key} went up: {baseline[key]} -> {now}")

# the ratchet: improvements become the new floor
save("gate_baseline.json", {k: min(baseline[k], current[k]) for k in current})

When an agent cleans something up, the bar moves with it, and no later change can quietly give the improvement back.

Agents are very good at making tests pass, which cuts both ways. I read test diffs more carefully than code diffs. A test that got "fixed" by loosening its assertion is the most common way a regression tries to sneak in.

Agents never see real data

The platform holds things that shouldn't leave the building: what was spent, with which vendors, by whom. Agents never work against that. Agents never read those records. When an agent needs to look at data, it gets a scrubbed copy of the database, with amounts, vendor names and personal details replaced by plausible fakes, while the shape of the data stays intact: row counts, relationships, and the odd edge cases that years of real use leave behind. Test suites run against throwaway copies, and what comes back to the agent is pass or fail and counts.

That gives agents real problems to work on without any of the parts that would hurt if they ended up in a transcript. When I need to know something about the live data, I ask for counts and yes-or-no answers, never rows.

Review it like someone's trying to break it

When an agent says it's finished, I often start a second agent whose only job is to find what's wrong with the first one's work. It reads the diff, runs the suite and tries the edge cases the spec named, plus a few it didn't. It starts without the first agent's context, so it doesn't share the first agent's assumptions. It regularly finds real bugs, mostly states nobody thought about: an empty list, a cancelled order, a user without the right role.

Then I read the diff myself. The adversarial pass tells me where to look. It doesn't replace looking.

Look at the screen

Tests tell you a page works, not that it looks right. Every UI change comes back with screenshots at desktop and phone widths, in light and dark themes, captured by a headless browser, and I look at every one. A label overlapping a number, a table that pushes the whole page sideways on a phone, a button that vanishes in dark mode: none of those fail a test, and all of them take ten seconds to spot. This site was checked the same way.

Where agents go wrong

Agents fail in a few recognizable ways. Most of the workflow above exists because of them.

  • Confident wrong claims. "All tests pass," when it ran one suite. "This function isn't used anywhere," when a template calls it. The fix is to ask for evidence instead of conclusions: the command it ran and what came back.
  • Scope creep. Asked to fix one bug, an agent will sometimes also rename things, reformat a file and "simplify" a function nobody mentioned. Each change might be fine; together they make the diff unreviewable. Specs say what not to touch, and diffs that wander go back.
  • Stale assumptions. Agents reason from what they saw earlier, even after the world has changed. While launching this site, an agent reported that the new deployment was exposing stray files. It wasn't. The machine's DNS cache still pointed at the old host's parking page, which answers every path with a 200. Checking from a second vantage point, a different resolver, settled it in a minute.

An agent's description of what it did is a claim. The test run, the screenshots and the diff are the evidence.

What I still do by hand

The data model. Decisions about migrating years of history out of the old systems. Anything that touches production: deploys, restarts, backups. And talking to the people who use the platform every day, which is where most of the good specs come from in the first place.

Agents changed how much one person can build. They didn't change what it takes to build something people rely on: knowing the domain, deciding carefully, and checking the work.

If you're working on applied AI in aviation or industrial operations, I'd like to compare notes. There's more about the platform in the case study.