The agentic SDLC is the software development lifecycle with coding agents doing most of the typing. Planning, building, testing, review, deployment and maintenance all still happen. What changes is where the time goes, because writing the code has stopped being the expensive part.
Most of the guides on this are written by companies that sell one piece of it, and they tend to describe the lifecycle around wherever their product sits. The review tool says review is the bottleneck. The static analysis tool says verification is. They’re not wrong. To be honest, verification really is one of the bottlenecks now. But you can’t buy your way out of it with one tool, because the problem starts much earlier, with the spec.
I’ve had coding agents merging their own work to production on my own product since January, and doing heavy agent-driven work for clients alongside that. This is the loop as I actually run it.
Discovery and the spec
The AI doesn’t write code for a very long time. Most of my time goes into planning, going backwards and forwards with it before any code exists. I set it off to do the research, read the existing code and pull in whatever it needs, and then we have a discussion. You don’t want to be hand-holding it. You want to set it a task, let it come back, and then do the direction setting and the challenging in one go. If you’re not careful it’ll ask you a question every ten minutes, which is really annoying, so I batch it: it asks all of its clarifying questions at once and I spend twenty minutes or an hour answering them.
I answer almost all of them with voice typing. I say far more than I’d ever type, and more context makes for a better spec.
It needs the context a person would need, too. If you want it to decide whether a feature still matters, it needs usage data. It’s perfectly capable of working out which features customers actually use and which they don’t, but only if it can see that information.
The other big input is meetings. Almost no teams record their calls. They spend three hours discussing a problem, the only record is in people’s heads, and then somebody has to go and write the ticket. My planning meetings are always recorded and transcribed, and the transcript goes straight into the spec. I wrote more about why the spec is the work now.
Review before anything is built
Once there’s a plan, it goes past a set of reviews before anything is built. I use Garry Tan’s gstack for a lot of this. It’s open source, and it’s mostly just skills, so it’s easy to poach from. Its auto plan runs a CEO review, an engineering review, a design review and a developer experience review where it’s relevant, and it runs each of them on both Claude and Codex. When two different models agree on a problem, it’s probably a real problem.
The reviews push back properly. On one engagement we deliberately took a plan to move an API from GraphQL to REST through it, and both CEO reviews came back hard against it. Developers not liking GraphQL isn’t a good enough reason to spend weeks to months changing it, and the reviews said so. You can override them, but it forces you to justify the decision.
There’s also an office hours skill that runs like a startup office hours session. It doesn’t look at the work at all. It asks why the thing exists in the first place and whether you should be doing it, which are the hard questions people tend not to ask themselves.
On top of that I run my own all hands skill. I’ve built up agent personas from the C-level down to UX designers and the maintenance programmer who just wants nothing to change, and they all ask questions. Twenty rounds is normal. It’s slow, but it’s a lot cheaper than finding the gaps after the code exists.
Building
The build is the part that’s changed most, and it’s the part I spend the least time on. I run everything on an always-on Mac mini with about eight sessions going at once in tmux, each plan in its own worktree with its own code name. They come up and down depending on what I’m working on.
The way I think about it is that I’ve moved from being a developer to managing a team of developers, which I’ve done in real life, so I’m comfortable with it. The process is basically you talking to a developer. The difference is that the developer never gets tired and has no idea of consequences, so the environment has to supply the judgment it doesn’t have.
Verification
This is the one the vendor guides are right about. The agent does an 80% job and tells you it’s done. Rather than getting frustrated with it, you need a set of tests and a harness that it can’t get past until the job is actually finished.
I used to be compassionate about this, in the way you would be with developers who hate waiting for tests. Then I realised the agent doesn’t care, and I made the gauntlet as hard as I could. I’ve written about how I let coding agents merge to production: the long test runs, review from more than one model, and the hooks that redirect it when it reaches for the wrong thing. And separately about the ways agents go wrong and what catches each one.
The principle underneath is old. Anything you want to happen in a codebase has to be enforced in code: in CI, in commit hooks, wherever. A team agreement gets forgotten the first time somebody new joins, and an agent is somebody new every session.
Deployment
Shipped means deployed. An agent will happily open a pull request and decide its work is done. It isn’t done until the deploy on the merge commit is green, with the acceptance checks included. With a merge queue in front of main, that’s what makes it possible to hand over a batch of work in the evening and find it deployed in the morning.
Maintenance
This is the phase the guides mostly skip, and it’s where a lot of the value is. I have a couple of skills that run on a schedule. Hygiene spins up an agent team every day to audit the codebase for the things linting can’t catch. It checks what it finds against the existing tickets so it doesn’t duplicate them, and anything new becomes a ticket. I can ignore that ticket or close it, but I always know what state the repo is in. There’s one check that flags every file over 300 lines and reminds me about it every single day. I’ve been ignoring it. It’s horrible.
Pulse is the other one. It goes back over everything that’s been marked done and checks that it actually is, because agents over-claim.
Agents also change the economics of old code. A lot of teams have a system that everybody says can’t be upgraded. On one engagement, a framework upgrade that had always been the sticking point, because of dependencies, was shown to work in about a day. Replatforming is almost always a mistake. The better way is a strangler fig: carve the old system out a piece at a time and keep shipping. The traditional risk with a strangler fig is that you get 95% of the way and never finish the last bit. That’s much less of a problem when the agents can do the tail.
What hasn’t changed
Almost none of this is new. Small, tightly controlled changes, automated testing, test and behaviour driven development, continuous delivery: everything Martin Fowler and his peers were pushing still applies, and it matters more now, not less. Plenty of teams got away with skipping it when humans were writing every line. Agents stop letting you get away with it.
What’s different is that the bottlenecks move. They move into the spec at the front and the verification at the back, and the people who used to spend most of their day writing code now spend it on both. I wrote about what that means for an engineering team, including the bill.
If I were starting this on a team tomorrow, I’d do three things first. Record and transcribe every planning conversation. Make the build hard enough that nothing gets through it unfinished. And put an agent on daily hygiene, so you always know what state the codebase is in.
If you want help setting this up on your own repo, that’s what team enablement is: two to five days on your codebase with your engineers.