🇺🇦 Stand with Ukraine — how to help

writing / 2026

What an AI-native agency actually sells when the code is cheap

05·10·2026 · 6 min read

The build is the fast part now. When I’m working on something, I spend the majority of my time talking to the AI before the work ever starts. I’ve got a long process and a bunch of skills that ask questions, research, and ask more questions, and I’ll go around for a good two hours on a thing before I kick it off. The rest of my time is verification: asking it to verify by running browsers, verifying it myself, being the QA, getting it to write Playwright tests. The bit in the middle, where the code actually gets written, is where I spend the least time.

That matters if you’re an agency or consultancy calling yourself AI-native. Most of what gets written about AI-native agencies says they sell outcomes and run small senior teams, and that’s fine as far as it goes. What it doesn’t say is where the work went. AI hasn’t removed the slow parts of a project. It’s shoved the bottlenecks into different places: into the PRDs, into the verification and into the merge. If you’re still selling build hours, you’re selling the part that got cheap.

I’ve written up the whole loop in what the agentic SDLC actually looks like, so I won’t go through it again. This is about what that loop means for a firm that sells delivery to clients.

The AI’s “done” isn’t done

Yes, you can spit out pull requests and hundreds of thousands of lines. Then you get bottlenecked in places the AI can’t help you. One of the things that becomes really apparent is that it says “yeah, this is great”, and you go and look at it and it’s not really that great. You can send it around a couple more times and it’ll get there. But its assertion that it’s done is a place where humans are going to be needed for a while, along with writing the spec and orchestrating the work.

That’s the accountability a client is paying a firm for. The usual line is that an AI-native agency owns the outcome. Owning the outcome means somebody who knows what good looks like goes and checks, and doesn’t take the agent’s word for it. That’s real time from senior people, and it doesn’t shrink because the code arrived faster. The vendor surveys are picking it up too: a Kantata survey run by Censuswide (200 professional services leaders, and Kantata sells professional services software, so read it with that in mind) found 89% of them spend significant time validating AI outputs.

The merge is the other one, and it catches people out. I’ve had five to ten workstreams running at a time, and you end up with it thrashing, trying to merge all of the work in. In a regular team everybody’s spaced out: John’s merging his PR on Tuesday and your work isn’t finished until Wednesday. These streams are all trying to merge at the same time, and they break each other’s work, and whoever merges second has a whole bunch of rework to do. A firm running agents across several client streams has that problem on every engagement, and somebody has to own it.

Scope stopped being the problem

Estimating used to be: I know I’m going to spend a day writing that and a day testing it. Now I’ve still got to do some of the testing, but the AI does most everything else, potentially instantly. Stuff that would have taken a month takes a day, and then there’s stuff it doesn’t really help with at all.

What that does to scope is the interesting part. As long as it’s clearly defined, there really can’t be scope that’s too big anymore. It used to be that developer time was the big problem. Now the problem is clearly defining the problem: being able to describe everything that needs to happen. In January and February I built an entire property portal from scratch, around half a million lines of code in about six weeks. I’ve got a background in real estate portals. Scope is not the same problem.

On side projects it’s almost waterfall in some ways. You need a very strong brief for the whole thing, because the AI works through a brief so quickly it almost overwhelms your ability to supply it with work.

For an agency that flips where the expensive part of an engagement sits. The brief used to be the cheap bit at the front, a couple of workshops before the real work started. Now the brief is most of what you’re selling, and the people who can write one are the ones a client is actually paying for.

Who you put on the client

This is the uncomfortable part for a lot of firms. The vast majority of software engineers are ticket takers and deliverers. I’m generalising, but if anything isn’t defined in the ticket, they break. That’s exactly the role AI is taking away. A fully described ticket, the AI can deliver the vast majority of. The bottlenecks move into checking that the AI has actually done its job properly, and into planning and discovery.

And a large number of developers just don’t have any experience in discovery. They’re not product managers. They don’t think about the way customers are going to experience it, and they don’t explore it in all the different ways it could go.

One product person writing every ticket and every PRD for a whole team of developers is just not a ratio that works anymore, right? Product becomes the bottleneck, and pretty much everyone who’s a developer needs to take on some level of product ownership to smooth it out again.

For a firm that’s a staffing question before it’s a tooling question. The people you put on a client need to be able to sit with the client, work out what the work actually is, write it down well enough that an agent can run with it, and then tell when the result isn’t right. If your bench is people who wait for a fully specified ticket, they’re doing the part the AI now does.

Then there’s the client who’s found Claude Code and decides they don’t need you: “I can do it in such a short time now.” Yes, you can build some stuff, but you’re not an engineer and you’re going to drive it off a cliff. It’s only a matter of time until it drops a production database. It’s dropped mine. I had backups, and I’ve got layers of protection around it for exactly that reason. Being able to explain that difference is part of what you sell. If a client still doesn’t see it, get good at firing customers. Tell them you’re not a good match, offer a referral to somebody who might be, and don’t burn the bridge.

Where this leaves the firm

Engineers aren’t done. I firmly believe they’re still required, not least for the judgment and the verification. The code the AI writes, when it’s watched, is probably what you’d have written anyway, and there’s no way a human typing on a keyboard can compete with that. The job moves up into orchestration and taste. Taste is more than architecture. It’s the overall ethos of how the project looks and how it runs over a long period.

There’s honest counter-evidence too. METR ran a randomised trial in early 2025 and found experienced open-source developers were about 19% slower with AI tools on mature codebases they knew well. To me that’s a reason to take the checking seriously, because the speed isn’t automatic.

The pitch to a client used to be about how many people and how many weeks. Now it’s closer to: I can do it in two weeks or I can do it in two days, what’s your choice? They’ll pick two days every time. What they still need from a firm is somebody who’ll specify the work properly at the front and won’t accept “done” at the back.

The gates I use for that check on my own product are in how I let coding agents merge to production.