🇺🇦 Stand with Ukraine — how to help

writing / 2026

Putting coding agents through their paces

28·09·2026 · 5 min read

I run coding agents hard, on my own products and on client work, and part of that is deliberately looking for the ways they go wrong. They go wrong in the same handful of ways, and once you know what to look for, each one gets caught early and turns into a guardrail. I’ve believed for a long time that anything you want to happen has to be enforced in code somewhere. You can’t have a discussion as a team and agree to do something, because somebody new joins and doesn’t do it. That’s even more true with AI. If you don’t want it doing something, fail the build when it does.

These are the failure modes I look for, and what catches each one.

Reading stale code

Agents will happily reason about whatever copy of the code is in front of them. On one client engagement, a root-cause sweep after a bug bash read a clone of the main repo that was eight days old, and five of its headline root causes were already fixed or wrong. The verification pass caught all five. On my own product, the primary checkout once drifted onto a branch that had already been merged, and an agent reading it nearly redesigned a PR on premises that were no longer true.

My own code goes stale in a different way. I moved one of my products from one cloud provider to another, 500,000 lines in under a week, and the agent still finds references to the old provider in that codebase and brings them up. Old context is as dangerous as old code.

The guardrails: every agent fetches and fast-forwards every repo before it makes any claim about the code, records the SHA it read, and stops and tells me if it can’t fetch. A session-start hook checks where the primary checkout is. If it’s sitting on a clean, already-merged branch it goes straight back to main, and otherwise it warns loudly and says to verify against origin/main. The shell hook won’t let an agent switch that checkout to a feature branch at all, and tells it to make a worktree instead. Stale docs and agent memory get deleted aggressively, because the agent will keep arguing for the previous direction.

Fanning out too far

Multi-agent runs make it easy to throw agents at a problem. Tens are fine. A 76-agent run built a QA plan for a client and it was good work. The next day a research job planned one agent per claim per lens to verify six notes, which came to 693 agents. I stopped it. The notes plus one write-up of the mechanism did the job.

The other thing to watch is loops. One multi-agent build tracked finished work by object reference, and results come back as copies, so nothing ever came off the queue. Each slice was rebuilt about ten times, 511 agents where about 75 would have done, until it hit the usage limit, about ten hours in.

The guardrail: before a big run, the agent multiplies it out, claims times lenses, and restructures if that’s much over fifty. Work is keyed by id, and it logs counts per wave, so a loop shows up early.

Taking shortcuts with git

Given a git task that gets awkward, an agent takes the shortcut. Asked to reopen a PR whose base branch had gone, one opened a replacement and force-pushed the original branch instead of telling me it was blocked. Asked for a rebase that hit a conflict, another squashed 14 commits into one and force-pushed, which took away the history the reviewer needed. Doing it properly meant resolving four conflicts by hand. Not hard, just slower.

The one worth knowing about is git reset --hard. An agent decided two newer PRs already contained all the work on an older one, pointed its branch at one of their tips and force-pushed. I looked at it and the two PRs didn’t add up to all the work on the original. Three audit passes accounted for everything, including 245 lines of tests, one of them the only regression test guarding a platform contract. They were restored.

The rules: reopen means gh pr reopen on the same number, and if that’s blocked the agent tells me and waits. Rebase means commit by commit with rerere on, every conflict by hand, then git range-diff and an empty diff against the expected tree. A backup branch is pushed first, --force-with-lease names the exact remote SHA, and nothing the task didn’t touch gets force-pushed.

The hooks do the enforcing. A pre-push hook checks every push that isn’t a fast-forward: a force-push to a protected branch, to anything already merged into main, or to a branch with someone else’s commits on it is blocked, and the message tells the agent what to do instead, open a PR or cut a new branch. A force-push to its own unmerged branch goes through with a nudge to use --force-with-lease. Commits and code edits on main are blocked and redirected to a worktree. When an agent tries to finish with uncommitted changes, the stop hook tells it to commit its own work and never discard anything it didn’t write. The hooks matter more than the instructions. Its default after a mistake is still to close the PR and open a new one, even with that written in AGENTS.md, and that’s exactly why the rules that matter live in hooks.

Doing CI’s job locally

Left alone, subagents re-run everything CI already runs. On a client’s legacy upgrade they were spending 18-plus minutes running the full unit suite twice and ten minutes booting against an empty database, and one left an app running for two hours and 47 minutes on a shared machine.

The guardrail: compile, run the specs for what it touched and the checks CI can’t do, push, gate on CI, and stop every process it starts.

Saying it’s done when it isn’t

Agents over-claim. The classic pattern is a chain of fixes: fix, fix again, fix again, each one declared done, none of them with a test. The hooks count them. A second fix to the same area on a branch gets tagged [self-correct] in the commit message automatically, two fixes produce a warning, and at three with no test the push is blocked until there’s a test covering the behaviour it keeps fixing. That one hook turns an agent going around in circles into a regression test.

They also write tests with escape hatches, where the test finds nothing and exits happily, and ask them directly and they’ll agree it’s a terrible test. They’ll still write the next one the same way. What I put in for teams is a specific review of tests in CI that fails on a bad test, plus a non-AI lint pass for the escape hatches you can detect by pattern.

The same goes for “FIXED” in an audit. An automated audit on my own site marked a hole in my game’s hall of fame as fixed, and it wasn’t. I wrote that one up in how I audit a codebase with AI. An agent’s “FIXED” is a claim to check, not a result.

Hooks instead of instructions

Wherever I can, a rule becomes a hook rather than another line in AGENTS.md, and the hook redirects rather than just saying no. Beyond the ones above: write SQL against a real database is blocked and pointed at migrations, committed migrations can’t be edited, infrastructure can’t be destroyed, a new third-party job or workflow service needs my approval, and editing a CI workflow gets a checklist of what to verify. How I set those up is in how I let coding agents merge to production.

Almost all of this is the old continuous delivery discipline. Agents just stop letting you get away with the gaps.