🇺🇦 Stand with Ukraine — how to help

writing / 2026

What changes for an engineering team when agents write the code

28·09·2026 · 6 min read

One of the first things I built at realestate.com.au was an agent website builder. It was a pretty complex Ruby on Rails application and I was very proud of it at the time. The product ended up going away, but I know how long it took to build the first version of it, which was about six, maybe eight months, with a team of eight developers and adjacent people around them.

Early this year I replicated the majority of it in about two days, three or four at most, with Claude Code and agent teams. I built the original, so I knew exactly what to spec, and it’s a first pass I’m still refining. Even so, it’s just undeniable that software development will be different forever. But the work doesn’t go away, right? The actual coding time is greatly compressed, and the bottlenecks just end up in different places.

Where the bottleneck went

A previous colleague was saying the only really important thing now is the spec, the decisions. If we took the spec, deleted everything and told it to put it back, it’d put it back in an hour. The hard thing is not the building now. It’s making the bloody decisions, and the merging, and verifying that it’s built it correctly.

What becomes really apparent is that the AI says, yeah, this is great, and then you go and look at it and it’s not really that great. You can send it around a couple more times and it’ll do the job, but its assertion that it’s done is where humans are going to be for a while. Rather than getting frustrated, it’s here’s a set of tests and here’s our harness, and you enhance the guardrails every time you see it going wrong.

And there’s bottlenecks in places like pull requests. I’ve had five to ten work streams running at any given time. In regular teams that’s usually not a problem, because everybody’s spaced out. Somebody’s merging their PR on Tuesday and your work’s not actually finished until Wednesday. No, all of these streams are trying to merge at the same time, and when they merge on each other they’ve broken each other’s work. Whoever merges second has to do a whole bunch of rework. I created a merge queue and told the agents in the CLAUDE.md that deployment was the target, not the PR.

The spec is the work now

By the time you’ve got a ticket to the point that it’s ready to estimate, it’s done. There’s an enormous amount of work getting it to that point, and that work is the specification. Estimating something as multiple days when the AI is going to pump out the vast majority of the code in 10 minutes, that’s a waste of time. There’s more to estimates than the code, I know, but that was the bulk of it.

On one engagement I just asked the AI to estimate the same tickets beside the team and compared how often it was right. The vast majority of the time its estimates were bang on. Let the AI do the estimates, and put the grooming time into the spec.

My process has been, the AI doesn’t write code for a very long time. I spend hours going backwards and forwards with it in planning, and my planning meetings are always recorded, always transcribed, because you feed everything back in.

Almost no teams record their calls. They spend three hours discussing a problem in grooming or design review, the only record is in their heads, and then somebody has to go and write the ticket. Take the transcript and have the AI turn it into tickets. Developers are going to hate this, because they hate meetings, right? But no, we should be spending whole days sitting there discussing stuff in infinite detail and recording every piece of it. And then all of a sudden the AI can pull complete epics out of those discussions.

I’m vastly generalising, but most software engineers now are ticket takers and deliverers, and the fully described ticket is exactly the part the AI can deliver. It can’t fall to product to write all the PRDs either, because then product just becomes the bottleneck. Pretty much everyone that’s a developer needs to take on some level of product ownership. Anyone in engineering with product savvy should be retraining into semi-product, because you still have to drive it heavily, and you have to know the customers as well as the architecture.

Any one piece of code you don’t have to live with

A lot of our heuristics have been around, I’m going to have to live with this for 10 years. If I don’t build it properly, the next person’s going to hate me, and there’s nothing so permanent as a temporary fix. But now with a spec, any one piece of code you don’t have to live with. On my property portal I want running costs as low as possible, so I converted the services to Rust, basically just going convert that one, convert that one. Memory on Lambda went from a gigabyte to 256 MB, and it’s faster. As long as the interface is fully defined, you can pull a piece out and put another piece in.

If you can replace any given piece of code that fast, then the quality of that code in isolation isn’t as important. The integrations and the specs and the testing around it are what’s important. On the portal there’s contract testing in the monorepo, so if the agent changes something in the GraphQL API that would break a mobile app, the build fails.

But you have to be super, super disciplined about tech debt, because it copies. It copies bad patterns just as much as good patterns. An engineer can work around and skip tech debt, but the AI just replicates it, and so it expands. And a lot of what people ascribe to intelligence in these models looks more like persistence. They keep building and building and building, and unless it’s tightly controlled you get more and more complicated code, until the code base is so big the AI can no longer reason about it.

The reality is almost all of the answers are in how we dealt with bad engineering before: continuous delivery, small tightly controlled changes, test and behaviour driven development. Almost no company actually did continuous delivery at the level the original book suggested, and it really needs to be that now. Anything you want to happen has to be enforced in code, so if you don’t want it to do something, fail the build if it does it.

The bill

I run two Claude Max 20x plans and a Codex Pro plan, about €200 a month each, and I max all three out. Used like that, each one is worth somewhere between 5 and 14 thousand dollars a month at API prices, depending on the plan, and I’m getting a return on it. Multiply that across a team at API rates, though, and it’s more than the cost of the team. It’s all heavily subsidised at the moment, but even at 10 times the cost it’s still cheaper than an engineer.

What I’d do is decide what an acceptable daily rate is. Make the top models something you ask for rather than the default, because they’re an order of magnitude more expensive, and train people to drop to lower models when they can. Take the average daily usage with the outliers removed. You still pay the bill, but that’s the number you plan on. And decide up front what happens if the budget blows mid-month, or people doing good work end up twiddling their thumbs without AI.

Engineers aren’t done

I started with assembly language, BASIC and C, and people were certain with all of those that the job would go. It doesn’t. The job changes. At one large engineering org, the objection that really got me was, oh, my skills are going to atrophy. Those skills are somewhat irrelevant now. They had these things where it’s like, oh, we cannot do that, that’s going to take three months of work. And then somebody went and basically did it on the weekend with AI, and all of a sudden it’s, okay, I need to change my prejudices a little bit. Pick one thing your team has written off as months of work and let someone have a go at it with AI.

And if you fire your engineers, they’re going to do something. If they can’t find other jobs they’ll start companies, and the barrier to entry that was previously employees is disappearing. The companies that fire their engineers are going to find themselves with a hundred new competitors.