writing / 2026
Writing specs for AI agents starts with a recorded PRD walkthrough
05·10·2026 · 7 min read
Most of what’s written about writing specs for AI agents is written for engineers. Addy Osmani, Allstacks and Augment have all put out good guides: templates, boundaries, what goes in the AGENTS.md file. It’s useful, but it starts at the point where somebody already knows what to build, and that part is the product manager’s job.
Give the AI a fully described ticket and it can deliver the vast majority of it. The bottlenecks move. They move into verifying that the AI has actually done its job properly, and into planning and discovery, which is where the spec comes from. This piece is about that part.
Detailed PRDs and brutal read-throughs
I did a bunch of work with a startup where the CEO was really adamant about super detailed PRDs before anything started. Then there was a read-through, where you had to explain in a lot of detail what you’d covered in research and why this was the approach that should be taken.
The two-weekly showcases were the same. He was not a mean person, but his questions were brutal. If you weren’t on top of your work and owning it, you were going to walk away from that feeling less than ideal. And every single one of those people stepped up their game. All of a sudden they were absolutely owning everything. It was very clear and very owned: this is what we’re doing, and this is how we’re doing it. That’s the way it should be.
That kind of detail matters more with agents in the loop, because the fully described ticket is exactly the thing they’re good at delivering. The detail has to come from somewhere, and that’s the PRD.
The walkthrough should be somewhat combative, and the PM still decides
A PRD walkthrough isn’t supposed to be “I have everything sorted out”. The point is to draw on everybody’s knowledge. And it’s not about consensus building either, because at the end of the day you’re the owner and you need to make the decisions. You take people’s opinions on board, discard some of them, and go, no, we’re doing it this way.
Combative is a bit of a rude word, but the best walkthroughs are somewhat combative. You get a lot of engagement and a lot of people going, have you considered this particular problem? Somebody points out it doesn’t fit the roadmap, because another piece is moving at the same time and there’s going to be a bunch of rework. That’s exactly what you want to hear before anyone builds anything.
Even a verbal PRD does this. Somebody starts describing a feature and I’m thinking, that’s pretty simple, that won’t take long. As soon as they start talking it through, there’s all of this nuance I hadn’t considered. Historically that’s why we do refinement at all, because so often I’ll be thinking about A, B and C, but not D and E. Another developer has that information: have you considered the queuing, have you considered the SLAs?
Bring something visual, too. People don’t read things, right? As much as the PRD is important, people are mostly visual. A throwaway prototype, or three or four different versions of the tricky bits, gives the room something to argue with that a document doesn’t.
Record it and let the transcript build the spec
Almost no teams record their calls. They spend hours discussing a problem in grooming or solution review, and they never record it. The only record is in their heads, and then somebody has to go and try to write a ticket from memory. Instead you could go: we’ve just spent three hours discussing this, all of the content is there, AI, take the transcript and turn it into tickets.
From a transcript you get a huge amount to feed back into the AI to refine the PRD. You shouldn’t have to be writing it all. Have a folder full of transcripts of discussions, and all of that becomes a big part of the spec building. That’s how I work on my own: most of my time is spent in a loop of asking the AI to do research, then arguing with it and correcting it. With a walkthrough, the recording makes your life so much easier, because you come out with this whole bulk of opinions on the other side.
Out of enough of those discussions the AI can pull complete epics, and even roadmaps. Developers are going to hate this, because they hate meetings. I’d still argue we should be spending whole days sitting there discussing things in a lot of detail, and recording every piece of it, and then letting the AI do the writing up.
For existing systems, let the agent write the docs, then go looking for lost requirements
Not everything starts with a blank page. Claude Code, Cursor and Codex will all go off, explore a code base and write the technical documentation for you. You do need to check it. I had this multiple times over the last year with one client, where it wrote exceptional technical documentation for legacy code.
The catch is the lost requirement. Some product manager asked for a feature years ago, it got built, and everybody has forgotten it exists. Ask an agent to move that system to a new platform and it will faithfully rebuild the feature nobody remembers. Whether that feature should still exist is a product call, and the generated docs are where you’ll spot it.
The same analysis works at the product level. If you can produce the technical spec, a PRD and some automated analysis of a product, you can use that as the case for shutting it down. I wrote up how I run that kind of analysis in how I audit a codebase with AI.
Give the PRD a deadline, and stop estimating what’s already specified
Put a deadline on the PRD. PRDs are as long as a piece of string, and you can keep building them and building them. The goal isn’t for it to be finished. It’s a working, living document. Tie the deadline to the walkthrough and get it in front of people.
The other thing that has to change is estimation. At one client a framework upgrade everyone had pegged at a month was done in two days once we actually tried it, and that’s a sign everything needs reassessing.
Even the process of estimating tickets is almost wasteful now. By the time you’ve got a ticket to the point where it’s ready to estimate, it’s done. There’s an enormous amount of work getting a ticket to that point, and that work is the specification. What’s left is verification. On one engagement I tried to get a team to use AI estimates, and they said no, we have to estimate. Instead I sat beside them, asked the AI to estimate the same work, and compared. The vast majority of the time its estimates were bang on, and at times it did a better job of analysing all the pieces than the engineers did themselves.
The research backs up how much the spec matters. GitHub looked at more than 2,500 agents.md files and found that most of them fail because they’re too vague. Osmani also points to a study that calls it the “curse of instructions”: models get worse as you pile more requirements on at once. A spec should be specific, and it shouldn’t try to carry everything in one go.
That fits how I’d run it anyway. You don’t need a perfect PRD with a perfect set of requirements. If what comes out is 90% there, that’s great, and you might even release it to customers. Then you write the next PRD, the set of features that closes the gap, and walk it through, record it and argue about it the same way.
The engineering guides cover what happens once the spec exists. Getting it to exist is still product’s job. If you want to see where this fits in the wider delivery loop, I wrote about what the agentic SDLC actually looks like.