🇺🇦 Stand with Ukraine — how to help

writing / 2026

What AI agents actually cost to run, and why the subscription price won't last

05·10·2026 · 6 min read

Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, and escalating costs is one of the reasons it gives, alongside unclear business value and weak risk controls. What I almost never see next to a number like that is what an agent actually costs to run. I’ve got some of those numbers from my own work, and they’re worth knowing before you plan a budget around agents, or plan to cut one.

Hundreds of dollars for one night of agents

The broad-scale changes I’ve been doing with agents mean running multiple agents at a time. One night earlier this year I had them going for five hours. I don’t pay per token, I’m on a subscription, but the tool calculates what the run would have cost at API prices, and it came to somewhere around five to six hundred dollars.

That was on the most expensive model. Even on a model at half the price you’re talking about two hundred and fifty dollars for five hours, for one small task. Scale that up to rewriting a large legacy platform and the token bill on its own gets big very quickly, and that’s before you know whether the approach even works on code in that state.

The subscription price is not the real price

All of the subscriptions, which I benefit from a lot, are massively subsidised, and that just can’t last. My suspicion is that the API prices reflect what it actually costs to run a model that size, at breakeven, not even a profit. The subscriptions sit well under that.

The cheaper models muddy it further. Some of them are quite capable, but it’s never clear how much they’re being subsidised either, so a low price today doesn’t tell you much about the price next year.

This matters most where a company is under pressure to cut costs and is pushing hard on AI with a best-case picture of what it can do. Aside from the capability gap, which exists, it’s not going to be as cheap as they think. The smaller models aren’t capable of managing that kind of complexity, and the token costs are going to go up as well, not down.

Copying the harness won’t give you the same result

There’s a related assumption I push back on. A leader sees what I’ve done with agents and expects to apply the same harness to their own code and get the same boost. I have a huge amount of concern about that.

Part of it is capability. The harness is only one piece of why those runs worked. The rest is deep, deep knowledge of the system and of the domain, and a capable model behind it. Part of it is cost. I can run agents for five hours because the subscription absorbs it. A team paying API rates for the same runs is looking at a vastly different bill, and the capability gap and the cost gap stack on top of each other.

Budgeting: per-token billing is hard to plan

On one client team I recommended moving to Claude Code, and part of the reason was the shape of the bill. When it was API-based, token-based, that was a problem, because it’s hard to budget for. Now that you can put people on a team plan with a monthly fee, that problem mostly goes away.

Another company I’ve worked inside handles it a different way. Teams already have model access through their cloud provider, so they aren’t that cost sensitive about internal experiments, and they’ve been fine with people running them.

If I were sitting with a finance team, I’d want both of those things. A flat, predictable line for the day-to-day seats, and visibility of what the same usage would cost at API rates. The tools already calculate that second number, and it’s the one to plan around, because the flat price is the subsidised one.

Spend tokens where they buy finished work

None of this is an argument for being stingy. We need to find ways of being very aggressive, and that means soaking up some of these token costs at times. But you have to do it in a way where you’re certain you’re not wasting money.

The waste is an agent working from half the picture. If you’re going to ask it to do a whole feature, it needs to have all of the context it could possibly need to do a 95% complete version, so you end up in a state where it’s done and you’re into tweaking and fine-tuning.

Once you’re there, the maths on rework changes. I think we should stop avoiding rework the way we used to. If it doesn’t come out right the first time, I just prompt the AI to adjust it. Having a perfect PRD with a perfect set of requirements is less important than it was. What came out was 90% there, great, you can even release that to customers, and the next PRD closes the gap. The tokens go into getting most of the way there in one pass, and the iterating after that is cheap.

A cheap weekly eval tells you if it’s getting better

The last piece is knowing whether the money is buying anything. For a conversational agent on one engagement I built an eval tool that runs a fixed set of questions against our agent and against other models, and scores the answers. It costs about five dollars in tokens to run all of the questions. At that price you can run it every week and actually graph how it changes. As we add more and more functions it should get better and better, and you can see that visually.

It can check the tool calls too. For a subset of questions it knows which tools should be called, so you can see the agent is still calling the right ones as more tools get added, and that new functionality isn’t hurting quality. And it finds problems. Early runs turned up questions where the service would just fail, which went straight back to the team, and one place where the evaluator itself needed fixing because it hadn’t been given the date.

What you end up with is really concrete numbers. There’s none of this, I asked ChatGPT one question and it was better. Across four or five hundred questions you know what it’s considering better and what it’s considering worse. Five dollars a week is nothing next to a single night of agents, and it’s what tells you whether the rest of the spend is working.

I’ve written about the other side of this, the ways agents go wrong and the guardrails that catch them, in putting coding agents through their paces.