Ask your AI assistant the same question twice and you may get two different answers. That is the technology working as designed. It also means the two answers cost different amounts, because you are billed on tokens and the second answer used a different number of them.
Scale that from one question to an institution and the annual AI line in your budget stops being a number. It becomes an estimate with no reliable error bar.
AI pricing moves in three directions at once. The price per token is falling, the volume of tokens is rising faster, and the tokens consumed by any single task vary run to run. That third one comes from how the models generate output, so no amount of usage policy removes it. The practical result is that you cannot hand a board a twelve-month AI figure and defend it, and the software vendors who embed AI in their products are finding they cannot quote you a stable price either.
Why can't anyone predict an AI bill?
Because the thing being metered is not stable. Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, put it in one line to the BBC: "it's a non-deterministic output, so it's a non-deterministic value."
That is a stronger claim than "usage varies." Usage varying is a management problem, and management problems have management fixes. This is different. Feed a model the identical prompt on Tuesday and Thursday and it can return a longer answer on Thursday, consume more tokens, and cost more. Nothing about your policy changed. Nothing about your headcount changed. The bill changed.
Agentic workloads multiply the effect. When several agents pass work between each other, each hop generates tokens, and a task that took four steps last week can take seven this week because an intermediate answer came back differently. Venters makes a related point about how easy the growth is to trigger: adding more AI agents is a click, where adding people involves a hiring conversation.
What are the three numbers that keep moving?
Your invoice is the product of three quantities. Goldman Sachs Research pinned down two of them in its May 2026 report Decoding the Agentic Economy, and the third comes from the models themselves.
| What moves | Which way | Who controls it |
|---|---|---|
| Price per token | Down 60 to 70 percent a year for inference | The model provider |
| Tokens consumed across the market | Up 24 times between 2026 and 2030, to 120 quadrillion a month | Nobody in particular |
| Tokens consumed by one task | Varies between runs of the same request | Nobody |
Falling unit prices sound like good news, and for a finance team they are a complication. A cost that drops 60 percent a year while volume climbs 24 times is not obviously going up or down. It depends entirely on where your institution's adoption curve meets the price curve, and you will find out afterward.
The public examples are instructive because they involve companies with unusually good cost discipline. Microsoft has reportedly pulled back its engineers' use of some third-party coding tools, and Uber is reported to have burned through a year of AI coding budget in a matter of months. Both accounts are hedged in the original reporting, so treat them as directional rather than precise. The direction is the point.
Doesn't consolidating vendors fix this?
Partly, and it is still the right move. We have argued elsewhere that tool sprawl multiplies the problem, citing research covered by CIO that found 85 percent of organizations miss their AI cost forecasts by more than 10 percent, with nearly a quarter missing by 50 percent or more. Fewer tools means fewer meters, fewer contracts, and fewer renewal dates.
Consolidation does not touch the underlying variance, though, and it would be dishonest of us to imply otherwise. Reduce fourteen tools to one and you still have a meter whose reading depends on how a model happened to answer. You have made the problem smaller and easier to see, without making it forecastable.
The same applies to prompt discipline. Rob Steele, CFO at the accounting software firm iplicit, told the BBC that companies need to be far more precise with prompts, comparing a vague prompt to sending a family member out for the weekly shop with no list. That is sound advice, and it narrows the range without producing a fixed number.
Why does this reach you through software you already bought?
This has changed recently, and it matters more for a regulated institution than the raw cost does.
When a software vendor builds AI features into a product you already license, that vendor now has a variable input cost. They are buying tokens too. So they face the same forecasting problem you do, one layer up, and they have to decide what to charge you before they know what it will cost them.
They are candid about not having solved it. Bill Peterson, senior director of product marketing at Sumo Logic, described his firm's internal discussions about pricing new agentic security services to the BBC: "Nobody's really figured it out." On what happens when model providers change their own pricing, he added: "You get into variable pricing, and it's changing every couple of months. Customers don't like that. That's not how anybody builds a budget."
Simon Gooch at the identity management firm Saviynt was blunter about multi-year commitments: "Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know."
A vendor saying they cannot commit to a three-year price is being honest. It is also a problem for an institution whose budget cycle is annual, whose board approves that budget, and whose third-party risk program assumes a contract with a knowable value. Your vendor management process was built for suppliers who can quote a price.
What should you require from a vendor who can't quote a stable price?
You are unlikely to get a fixed price from a vendor reselling someone else's tokens. You can still contract like an institution that expects to be audited.
- A ceiling, not only a rate. A per-token or per-credit rate tells you the slope. It does not tell you the maximum. Ask what the annual figure cannot exceed, and what happens when it is reached.
- A notice period before pricing changes. When Anthropic moved agent workloads onto metered credits, customers got about 30 days. We wrote about what that shift means for institutional workloads. Thirty days is not long enough to rebuild a budget or migrate a workflow.
- Access to your own consumption data. If you cannot see token usage by team, by workflow, and over time, you cannot forecast next year or explain last year. Ask whether that data is available, in what format, and how far back.
- A written answer on renewal. What happens if the vendor's input costs rise between now and renewal? Who absorbs it? Get the answer in the contract rather than in an email from a sales engineer.
- Clarity on who bears the variance. Per-seat pricing puts it on the vendor. Per-token pricing puts it on you. Outcome-based pricing splits it in a way that depends heavily on the definition of the outcome. All three are legitimate. Only one of them lets you write a number in a budget and be right.
What changes with a fixed cost?
Running the model on your own infrastructure changes what kind of number the AI line is. The cost becomes the hardware, the software, and the people who run it. Usage rises and the number stays where it is. A department triples its queries and the number stays where it is. That is the whole claim.
We are not going to tell you it is always cheaper, because that depends on your volume and we have laid out the arithmetic elsewhere. At low volume, metered cloud AI is often the better economic call, and any honest comparison starts there. What a fixed cost buys, at any volume, is a figure you can put in a board packet twelve months out and defend when someone asks how you arrived at it.
There is a real trade, and you should price it in. You take on capacity planning. If you size the hardware for today and your usage triples, you buy more hardware, and that is a procurement cycle rather than a line item that absorbs the growth on its own. The variance moves from your invoice to your infrastructure plan. For most regulated institutions that is the better place for it, because infrastructure plans are annual and boards already understand them.
Cognetryx runs inside your network on your hardware, which is why the pricing is a subscription rather than a meter. You can read how the four-year comparison works out against per-token cloud billing.
The question to take into your next renewal
Every AI vendor conversation for the next two years will include a pricing slide. The useful question is not what the rate is. It is what the rate does when the vendor's own costs move, and who signed up to absorb that.
If the vendor cannot answer, they have not decided yet, which is a fair position for them to be in this year. Just be clear that undecided means the risk defaults to whoever has less bargaining power at renewal.
See what a fixed AI cost looks like
A short assessment shows what private AI costs to run at your volume, as one number that does not move with usage.
Book a Free AI Strategy Assessment →Sources: Joe Fay, "Tokenomics: Why making AI pay is tricky," BBC News, August 2026, for the quotations from Will Venters (London School of Economics), Bill Peterson (Sumo Logic), Simon Gooch (Saviynt), Rob Steele (iplicit), and Oliver King-Smith (smartR AI), and for the reported Microsoft and Uber examples, both of which the BBC hedged. Goldman Sachs Research, Decoding the Agentic Economy, May 2026, for the forecast of a 24-fold rise in token consumption to 120 quadrillion tokens per month between 2026 and 2030, and for inference cost per token falling 60 to 70 percent per year. Forecast-accuracy figures via CIO, as cited in our analysis of AI tool sprawl. This article is informational and not financial or procurement advice.
Frequently asked questions
Why are AI costs so hard to predict?
Because the same request does not consume the same number of tokens twice. Language models produce non-deterministic output, so an identical prompt can return a longer or shorter answer and bill differently. Two other numbers move at the same time. Goldman Sachs Research expects token consumption to rise 24 times between 2026 and 2030 while inference cost per token falls 60 to 70 percent a year. Your invoice is the product of three quantities, and none of them holds still.
Does consolidating AI vendors make costs predictable?
It reduces how many meters you are reading, which is worth doing, and it does not make any single meter predictable. Consolidation addresses sprawl. It does not address non-determinism, because the variance comes from how the model generates output rather than from how many tools you bought. A single well-governed tool still bills differently for the same task on different days.
Why can't software vendors quote a stable AI price?
Because their own input cost is variable. When a vendor builds AI features on someone else's model, the model provider's pricing and the token consumption of each customer both move underneath them. Vendors are openly saying they have not solved it. Bill Peterson of Sumo Logic told the BBC that nobody has really figured out how to charge for agentic services, and that variable pricing changing every couple of months is not how anybody builds a budget.
How should you contract for AI when the price keeps moving?
Ask for a ceiling rather than only a rate, a notice period before pricing changes, access to your own consumption data, and a written answer on what happens at renewal if the vendor's input costs move. Establish who absorbs the variance. If the answer is that you do, then the vendor has moved an unforecastable cost onto your budget and called it a subscription.