🇺🇦 Stand with Ukraine — how to help

writing / 2026

How I audit a codebase with AI, and what it found on my own site

28·09·2026 · 6 min read

When I come to a repository I don’t know, the first thing I do now is ask the AI to write a technical spec, where it tries to analyse the entire code base, and a product requirements document. As in, reverse-engineer what the thing is actually trying to do. It’s phenomenal for onboarding to a service that isn’t well documented, and it’s where most of my audits start now.

A bunch of what it writes is going to be wrong. The next step is reading it with the people who know the system and going: no, that’s wrong, no, that’s wrong, it’s lacking context. And when it’s missing context, working out why. Either it’s in somebody’s head and not written down anywhere, or it’s in a system the AI doesn’t have access to.

And sometimes it’s the other way round. I’ve had multiple examples over the last year where we pointed it at a legacy code base, it said something was the case, and the first reaction in the room was, no, the AI’s just wrong, it’s hallucinating. Then somebody goes off and looks, and yep, it’s there. Some product manager asked for it years ago and nobody remembers. If you try and move to something new straight off the code, it rebuilds a feature that everybody’s forgotten exists.

Earlier this year I took a very large, very legacy code base with a specification template, and it split up into hundreds of sub-agents, one per external path, tracing each one through. What that doesn’t give you is whether those paths are in use. When did somebody last visit that page or call that endpoint? In legacy code bases there’s a huge amount of stuff we don’t have to port forward.

I give the AI read-only access to pretty much everything I can: analytics, observability, all of it. The more it can see, the better its analysis, and it can pull out audience numbers that tell you which parts actually matter. On my own site that was just a little command-line tool over the analytics, so the agent could pull page views itself.

How bad, and how sure

First I pick the two or three subsystems that would be catastrophic if they failed for what you’re actually trying to do, and most of the time goes there.

When I ran one over a big legacy code base, the defect ledger came back with a huge number of P0s, and I wouldn’t even show that screenshot until I’d validated them. Whether the AI has classified something as a P0 and whether I would is another matter. Say it flags a password committed to the repo. If it’s a dev password, theoretically at least it’s overridden in production. I’d be extremely surprised if it was the production one, but you can’t presume that, so I need to validate it. If it’s a real API key, it’s compromised and it needs rotating. But if it doesn’t have an immediate negative impact on customers, I wouldn’t call it a P0.

Every finding gets two things on it: how bad would it be if it’s true, and how sure am I? And I’m only really sure if I can point at the file and the line, and another engineer would look at it and go, yep, that’s right. If it’s catastrophic and I can’t prove it from the repo, I don’t put it in as a fact. It goes on a list of spikes, things somebody needs to go and check before anyone makes a decision off them. But I don’t downgrade it either. It stays catastrophic, it just sits on that list until somebody’s validated it.

Because the AI’s done this massive scan, it’s gone shallow rather than deep. Before I commit to anything I send off sub-tasks: take this one, go deep, is it actually accurate, and what’s the impact? To really verify the security ones you have to step through them and follow the instructions. I’ve done two or three, and when one didn’t work it was likely some other guard it wasn’t aware of. The code itself is absolutely vulnerable in the way it says, and something else is protecting you. I’ll also run the same audit with two different models, because they miss different things, and have one deduplicate against the other so we don’t end up with unnecessary tickets.

Pointing it at my own stuff

In February I had five specialist agents audit my own site, including the prompt-engineering game on it. They flagged that the hall of fame didn’t validate scores, clamped the score between 0 and 10,000, and marked it fixed. In May somebody posted 999 under the name “this is vulnerable?“. Which, yeah, it was. You can’t legitimately get anywhere near 999, and the entry had zero zones completed. The server was still believing whatever score the browser sent, as long as it was under 10,000. The same agents that did the fix had marked it fixed, and I never went back and checked.

In September I went back with a read-only pass that ranked everything it thought was nonsense by traffic times wrongness, and the forged score was top of the list. Now the server works the score out from graded attempts, a session can only claim a place once and only after clearing every scene, and there’s a test for forged, incomplete and replayed submissions. After the deploy we posted a 999 straight at the endpoint, the way the old code took it, and it came back 400.

The more embarrassing part was the content. A lot of the older site had been AI-written, and I hadn’t checked it properly. There was a “38% improvement” figure that came from nowhere. The prompt library said “30+” prompts and there’s 20 on the page. There was a blog post credited to somebody who doesn’t exist, and stuff about “our team”, when there isn’t one. Some of the research notes were literally tooling output. The cleanup deleted 611 files of slop, duplicates and dead code.

Tests and CI

A lot of teams want to know if they can point agents at their code and go fast. I’ve been able to rewrite things rapidly on my own platform because the code base actively supports agentic development. It’s well factored, it has a large number of tests, it can verify that the work is complete in a lot of cases, and the seams are clear, so any one piece can be pulled out and put back in. On a legacy estate with not much test coverage, a lot of dark code and unclear seams, you can still use agents, but you’re going to have to hand-hold them.

When I ran the full suite over my own platform, a lot of the findings were the same root cause filed once per lens, and a good chunk were strengths, not problems. But it did flag a dashboard query that didn’t match the schema in the repo. It couldn’t check that against the live system, so that part’s still to confirm. The end-to-end test that should have caught it was just matching a string, though, so CI would have stayed green either way.

When I look at tests and CI in an audit, I’m really asking: could an agent get something wrong through this? Can it tell from CI alone that the work is done? Do the end-to-end tests check what the user actually sees? How long does a human take to verify a PR? I can put out ten PRs, but if I can’t verify them I can’t merge them.

What I hand over

You get a verdict that answers the question you came with, like can we replatform this, or is it safe to point agents at it. Then the findings as tickets with their evidence, the spike list, and a list of what I didn’t probe. Anything that needs the live system goes on that list, and so do legal questions, which belong with a lawyer. It’s not a pen test either, where somebody’s actually breaking things. I compile it into a PDF with executive summaries so it’s more readable. I’ve written an assessment before that pretty much got put in a drawer, so I care more about the findings becoming tickets somebody owns than about the document.

If you want this done on your code base, that’s the two-week architecture audit.