🇺🇦 Stand with Ukraine — how to help

writing / 2026

How I let coding agents merge to production

28·09·2026 · 6 min read

Anything I want enforced with coding agents is enforced by failing the build, not in PR review. I’ve been doing this long enough that I predate pull requests, and I’ve always been a continuous delivery and automated testing person. With agents that matters more, because nobody is reading all of the code anymore.

Since January I’ve run my own property portal this way, and the risks in this post are deliberate. The portal is a project I use to push the boundaries on purpose: the agent gets more access and more autonomy than it ever would on client work, so I find out where it breaks. Client work is more complex, and it doesn’t run with this much freedom.

The portal is around half a million lines of code, built in about six weeks. The agent has its own machine, access to GitHub and the hosting, production credentials, and all of the logging and alerting. It works in worktrees, and when a PR gets through the build it merges itself and deploys to production.

The genius junior, and what’s changed

When I started this in January, the way I thought about it was a genius, extremely talented junior engineer who hadn’t got the battle scars. A senior has deleted production once and not slept for a month afterwards. The agent hadn’t, and it would happily do the thing a senior never would. Mine dropped a production database earlier this year. The backups made it a non-event, and the access it had is now behind the hooks further down.

That picture is changing. The newer models are noticeably more cautious, and more careful not to do things the same way. What hasn’t changed is that the agent has no sense of consequences. It has no stake in the outcome: when the session ends it effectively ceases to exist, whether it did a good job or not. And the models aren’t interchangeable. Opus 5.5, Fable and GPT-6 Astra all behave differently, so the guardrails you need depend on the model you’re running.

The old adage still applies: if a junior breaks production, that’s a team system problem, not the junior’s fault. On an older project that uses Prisma, the agent ran db push a couple of times, which changes the database directly instead of going through migrations. The answer isn’t a sterner prompt. I took the command away. Every mistake it makes points at a gap in the environment or the safeguards, and closing that gap is the job.

Make the gauntlet mean

With human teams you keep builds short, because developers hate waiting. Early this year I realised I was extending that courtesy to the agent, and the agent doesn’t care. Now I make it very hard for it to merge. I want so many tests that getting through takes 30 minutes, an hour, two hours, and I don’t mind if the end-to-end suite runs fifty times. The agent waits on each check in the background, and when one fails it goes around again. There are agents I don’t touch for half a day.

AI review goes in the gauntlet too. I’ve had Codex set up as a reviewer so strict that it said no, no and no again, and Claude fixed every problem before it finally merged. I’ve also used Greptile, a very aggressive pull request reviewer, with the rule that nothing merges until every comment is dealt with. If you want GitHub to enforce that part, turn on “require conversation resolution before merging”.

Watch the tests themselves. Agents are good at writing escape hatches, a test that finds nothing and passes anyway, and they’ll try to disown breaks they didn’t cause: that’s not in our code. It doesn’t matter whose code it is. Somebody has to fix the build, and the build is how you make sure somebody does.

Hooks that redirect

When I see the agent do the wrong thing, the first fix is a Claude Code hook. Hooks run before and after tool use, and when the agent stops. How you write them matters. Claude is smart enough to work around a flat block, and what it does instead is unpredictable. A good hook doesn’t just say no. It says use this instead.

It loves raw SQL, and given the chance it will mutate the database rather than write a migration. Select queries are fine, and a write gets stopped before the command runs:

echo "BLOCKED: Never run write SQL against production databases." >&2
echo "All data changes go through migration files. Read-only SELECT queries are fine." >&2
exit 2

Exit 2 blocks the command, and whatever you print is what Claude reads next. The migration hook works the same way: if a migration file is already committed, the agent can’t edit it, and the message tells it why and gives it the command to generate a new one.

Most of these hooks came from watching it repeat a mistake. Every so often I get the AI to go back through old sessions looking for things I’ve corrected, and fold them into the build and the docs. One pass in March turned up more than 22 incidents across my projects, and the recurring ones became hooks and lint checks. Some go stale as the models improve, and what one model needs another doesn’t, so they get reviewed whenever I change models.

Deployment is the target

Auto-merge without a queue thrashes: ten jobs hit the merge point at the same time and burn through GitHub Actions credit. I went through that, briefly ran my own CI server, and came back to GitHub Actions with a merge queue and one rule in the agent’s instructions: shipped means deployed. An agent will happily open a PR and decide its work is done. It isn’t done until the deploy on the merge commit is green, acceptance checks included. There are fourteen required checks before anything reaches main: lint, builds, typechecks, unit and database tests.

Merged still isn’t done. In September a performance change merged and every deploy job passed, and then the check at the end of the deploy measured cold search at nearly five seconds at p95, against a budget of three. The ticket stayed open.

That’s what lets me hand it a batch of work overnight. I gave it a brief at quarter past one in the morning, and by about nine twenty PRs had merged through the queue, with the last deploy green on every check. A couple in between went red on the search speed check. The only thing it left open was a rewrite of the vision doc, because that’s a founder decision and it waits for my sign-off.

What I’d do on a team

My portal runs a single environment to keep costs down, so the end-to-end checks run after the deploy. On a team I’d gate production behind a production-like environment. Differences between environments are a killer, and framework upgrades are where you see it: the build passes locally and breaks once it’s pushed, and that’s not the agent’s fault.

I’d also score risk rather than auto-merge everything. One team I know has the AI rate every PR from one to five on how dangerous it is, and anything rated a one or a two merges without an approval, so a one-line copy change isn’t waiting on anybody. They had humans rate the same PRs, and the AI was almost always accurate, and more conservative at times. Humans end up looking at the PRs that actually need them. Add feature toggles so new things go out switched off.

Give the agent access to everything a developer would look at. If it can’t reach something, getting it that access is your job. Database access should be read-only, because it will try to fix problems with raw SQL.

Everything in the Continuous Delivery book still applies: automated testing, guardrails, small changes that are tightly controlled, test and behaviour driven development. It matters more now, not less.

If you want help setting this up on your own repo, that’s what team enablement is: two to five days working on your codebase with your engineers.