Josh DayAgent Playbook

Their Word Doesn't Count

How I rebuilt all 53 screens of the Boost admin with a team of AI agents, and the system that decides when they're done.

I sent a coding agent to paint a sidebar purple. It came back with a timezone bug.

If you set up an A/B test in the Boost admin after 9pm Pacific, it failed with "start date in the past." It wasn't in the past. Nobody asked the agent to look at dates. It looked anyway, fixed it, and the fix passed every check we have.

That agent was one of a team that rebuilt the entire Boost admin, 53 screens, onto Shopify's new design system. Every screen builds. Every screen passes its tests. Every screen scores at least 7 out of 10 from a critic that never learns who built it. I wrote none of the code.

Getting agents to write code is the easy part. Anyone with a chat window can get a page back. The hard part is knowing what's true. An agent says "done" in the same voice whether it's done or not. On one page, you check. On 53, you can't, and if you ship on the agent's word, you're shipping a guess.

So here's the system I use to run a team of coding agents on a real product. Ten moves, in order, and you can copy most of them tomorrow. Where it stands, honestly: it's all on a branch. Nothing has shipped. I review before it does, and that review is move ten.

The cast

Think of it as renovating a house with the family still living in it. Someone decides what it should look like. Someone runs the job site. Crews swing the hammers.

I set the taste and hold the release key. I decide what good looks like, I break the tie when two good options disagree, and nothing ships without me.

Hermes runs the job site. It's my main agent, and it writes zero app code. It writes the plan, splits the work, records every ruling I make, and re-runs every check itself.

Prime agents are the crews. They write the code, each one in its own lane.

Big AI builds rarely break in the coding. They break in the managing. So give the managing to an agent whose only job is managing, and give it rules it can check.

1. Map it before anyone touches it

Before any agent wrote a line, we mapped the admin into a graph: every page, every component it uses, every API call it makes, every test that covers it.

An agent without a map guesses. It opens a file, follows an import, opens another, and builds its picture of the app one hallway at a time. That's how you get agents calling routes that don't exist and missing the test for the page they just broke.

With the graph, an agent handed "checkout upsells" knows on day one every file, call and test it owns. The graph also shows the parts almost every page leans on. Those are the load-bearing walls. They get rebuilt first.

2. Hand them today's manual

Agents write from memory, and their memory is out of date.

Polaris is Shopify's design system: the buttons, cards and layouts that make an app look like it belongs inside Shopify. When we started, Shopify's own docs tool only knew the previous version. Ask an agent for the new one and it confidently hands you the old one.

So we kept a local copy of the new Polaris docs, the icon list and the migration guide, and every agent worked from that copy instead of its memory.

Then we screenshotted Shopify's own new admin, and the agents measured spacing and type sizes straight off it. "Make it look like Shopify" is a vibe. A screenshot with measurements is a spec.

3. Write the rules down, then lock the doors

Hermes wrote the plan and recorded my rulings. Four of them:

  • Pure new Polaris, light, native to the new Shopify admin.
  • Keep the newer version of every page our team had already built: its structure, its data, its behavior.
  • Analytics come from the new data layer only, with no quiet fallback to the old one.
  • Nothing goes live. Everything stays on a branch.

The third one exists because agents love a fallback. When the new path gets hard, they quietly route around it to the old one. The screen looks right. The numbers come from the wrong place. You find out months later. A ruling on paper is one no agent can forget.

Paper only goes so far, though. Rules in a prompt are suggestions. The rules that mattered most were locks. The agents could not push code to GitHub, and they could not touch Firebase, where the live data sits. That wasn't a sentence asking nicely. It lived in the harness: the agents' git and firebase commands were wrapped so a push or a deploy simply failed. A confused agent could make a mess on the branch. It had no way to carry the mess out the door.

4. Prove it on three pages

Before scaling up, one agent rebuilt three pages: Moneyboard (Boost's results dashboard), Cart, and Rewards.

Three pages is small enough to check by hand and big enough to hit the real problems: the look, the data wiring, the process. If the approach was wrong, I wanted to find out on three pages, not fifty.

It worked. That's where the question stopped being "can agents do this?" and became "how do I run six at once?"

5. Build the shared parts once

The fastest way to get six different admins is to let six agents each build their own date picker.

So one agent built every shared piece first: product and collection pickers, confirm dialogs, the save bar, page headers, metric cards, charts, empty states, the command palette, the sidebar and the header. Then it wrote the guide to using them.

It also wrote a file map that gives every file in the admin to exactly one section. That map is what makes parallel work safe. No two agents ever edit the same file, so no agent quietly overwrites another's work.

6. Six agents, six lanes

Then six Prime agents started at once, one per section. Each got its own copy of the code and its own list of files it was allowed to touch.

Each one works in two phases. First it scouts: it reads every screen in its section, takes before screenshots, and writes a plan covering every screen, redirect, test and risk. It can't write code yet, and the check enforces that. Only after the plan is reviewed does it build, on the shared kit.

Scouting feels slow. It's the cheapest hour in the whole project.

7. Attack the plans before anyone codes

If I could keep only one move, it's this one.

A separate agent read all six plans side by side and checked every claim against the real code. It hunted for files two plans both claimed, shared parts nobody had built, API calls to routes that don't exist, and tests filed under the wrong section.

First pass: 8 real problems. One plan had a product picker calling a route that didn't exist. Others expected helper files that weren't there. Left alone, six agents would each have hit those mid-build and patched around them six different ways. We fixed them once, in the shared kit.

A mistake in a plan costs minutes. The same mistake with six agents building on top of it costs hours, and six different patches.

8. Their word doesn't count

Agents are optimistic about their own work. Every one of them will tell you it's done. So their word doesn't count. Done means passing a check the agent can't talk its way past:

  • The admin builds, and lint, type checks and unit tests show no new problems compared with the current version.
  • Every click-through test for the section passes, with none skipped. (These are scripts that click through the admin the way a person would.)
  • A blind critic scores a screenshot of every screen, and every screen needs 7 out of 10. The critic never knows which agent built it.

"None skipped" is there on purpose. The easiest way to pass a test is to turn it off.

Then Hermes re-runs the whole check itself before it believes anything. The agents aren't lying. But an agent that says "all tests pass" may have run the wrong tests, skipped the slow ones, or read an old log. The only report that counts is the one Hermes ran itself.

9. Merge, then check again

When the six sections were done, they merged into one, and the whole check ran again on the combined admin. Pieces that pass alone can still fight once they're together.

All 53 screens pass.

Then I made a taste call: Boost purple in the sidebar. One Prime run folded it in across the admin. The critic caught five visual bugs the change introduced, and they were fixed before I saw a single one.

That's the run where the timezone bug turned up. Prime was there to change a color. It noticed that A/B tests set up after 9pm Pacific failed with "start date in the past": the form read the date off your clock, the scheduler read it in Eastern time. It fixed it, and the fix passed the full check.

Nobody asked it to look. That's what you get when an agent has a map of the whole app, checks that can prove a fix, and room to notice.

10. I hold the key

It's still on a branch. I look at the before and after screenshots, and I click through it myself. Nothing reaches GitHub, Firebase or a live store until I make that call.

That review is the job, not a formality. The checks tell me the admin works. Only I can tell whether it's right: whether the page a merchant opens first is the page she needs, whether purple was the right call. The agents are the crews. The checks are the inspector. The owner still walks the house before the family moves back in.

What I'd copy for any big AI build

  1. Map the code before anyone touches it.
  2. Give agents today's docs, not their memory.
  3. Write your rulings down, and put your hardest rules in locks, not prompts.
  4. Pilot on a few pages before you scale.
  5. Build the shared parts once, then go parallel, with every file owned by one lane.
  6. Scout and plan before coding, and have a second agent attack the plan.
  7. Accept "done" only from a check the agent can't talk its way past, and re-run it yourself.
  8. Keep the release key.

The agents write the code. The check decides what's true. I decide what ships.