AI Strategy9 September 20267 min read

Why Do We Keep Paying for AI? — Part 3 of 8

Per-seat vs per-use: the two ways AI is priced

One AI product bills you per person. The other bills you per use, measured in tokens. Confuse the two in your budget and one of them will surprise you — usually in the month the service starts working.

#AI Strategy#Cloud Cost#FinOps#Leadership

The previous article drew the line between AI for your people and AI in your products. This one is about what happens to that line when it reaches the finance system, because the two sides bill on completely different principles.

One counts people. The other counts work. That difference is why one of them sits quietly in a forecast for a year and the other one produces a phone call.

What each invoice is actually countingPER-USER LICENCEpersonpersonpersonpersonpersonpersonpersonpersonEight seats, eight fees,the same number every monthPER-USE CONSUMPTIONA quiet monthA steady monthThe month it caught onSame price per unit. The bill follows the bar.A fixed cost per head and a variable cost per use are different animals.They belong in different budget lines and are governed differently.
The left-hand invoice counts people and does not care what they did. The right-hand invoice counts work and does not care who did it. Budget them in the same line and one of them will surprise you.

Per-user licensing: a fixed cost per head

This is the model most organisations already understand, because it is how they buy nearly all their other software.

You pay a fixed monthly fee for each person who has access. If the fee is R400 a month and you license 500 people, you pay R200,000 a month. If those 500 people have a spectacularly productive month and use the tool for six hours a day each, you pay R200,000. If half of them forget it exists, you also pay R200,000.

The strengths are obvious. It is predictable — it can be forecast a year out with confidence, because it moves only when headcount moves. It is familiar — it slots into existing software asset management, existing renewal cycles, existing approval routes. And it is capped — there is no scenario in which enthusiastic usage produces an invoice nobody expected.

The weakness is on the other side of exactly the same property: you pay for the people, not for the value. A licence assigned to somebody who opens it twice a quarter costs the same as one assigned to somebody who has restructured their working week around it. That is the quiet waste in nearly every seat-based estate, and it is invisible unless someone specifically looks at assignment against actual usage. It is also, happily, the easiest AI money any organisation can save.

Consumption pricing: a variable cost per use

The other model bills you for what the AI actually processes.

Nobody is licensed. There are no seats. There is a service running, and every time it does a piece of work — reads a question, considers some of your content, writes an answer — that work is measured and charged. A quiet month is cheap. A busy month is not.

This is the model that feels endless to leaders, and I understand why. There is no natural stopping point built into it, no moment where the thing is bought and paid for. It just keeps metering, the way the electricity meter in your building keeps metering.

But "endless" is the wrong word for it. The right word is proportional. The bill grows because the usage grows, and usage growing is, in almost every case, the thing you were hoping would happen.

Tokens, in plain language

The unit these services meter is the token, and it is worth thirty seconds of explanation because it turns up on every invoice and in every estimate.

A token is roughly a word — a bit less, actually; long or unusual words get split into two or three, and punctuation counts. For practical purposes, treat it as the words going in and the words coming out.

What you are actually paying for, per useONE INTERACTIONINThe question thecustomer typesINThe policy text youattach so it answers wellOUTThe answer itwrites backAll three are counted in tokens — roughly, the words going in and coming back outAND THEN MULTIPLYONE CONVERSATIONa few hundred words each wayA HUNDRED A DAYa busy help desk, quietlyTHIRTY THOUSAND A MONTHa service the public has foundThis is why it feels endless. It is not endless — it is proportional.
Nothing about the price changes between the first box and the last. The volume changes. Once that is clear, the consumption bill stops looking arbitrary and starts looking like a demand curve, which is something a leader already knows how to manage.

Three things get counted in a single interaction, and only the first is obvious.

What the user typed. The question itself. Usually short.

What you sent along with it. This is the part that surprises people. To answer well, the service typically attaches relevant material — the policy extract, the product details, the customer's own record, the instructions telling it how to behave. That supporting content is often far larger than the question, and it is counted too.

What the AI wrote back. The answer. Usually charged at a higher rate per token than the input, because generating is more expensive than reading.

Add those up and one interaction has a cost. It is a small cost. The reason the monthly figure is not small is arithmetic: a small cost multiplied by a number that grows every time the service gets more popular.

Why this matters more than it sounds

Understanding the unit changes what you can ask for, and that is the practical payoff.

It means you can ask what one interaction costs — a number your team can actually produce, and one that turns an abstract worry into a line of arithmetic you can do in your head. If a resolved query costs a few Rand and it replaces a call that costs considerably more to handle, you have a business case. If it costs more than the call, you have a problem worth naming early.

It means you can ask what is in the payload — because the supporting content attached to each question is a design decision, not a fact of nature. Sending an entire policy manual with every query, when a relevant page would do, is a real and common way to multiply a bill several times over for no gain in answer quality.

And it means you can ask what happens at ten times the volume, which is the question that separates a business case from a hope.

The budgeting mistake to avoid

Here is the thing that actually goes wrong, and it is a filing error rather than a technology one.

The two costs get merged into a single "AI" line in the budget. A number is agreed. Then the consumption side grows — because the service launched, or was promoted, or a competitor's phone line got worse — and it eats the room that had been assumed for the licence side. Or the reverse: a licence renewal lands, absorbs the line, and the customer service has to be throttled to fit.

They do not belong together. One is a fixed cost driven by headcount, reviewed annually, owned by whoever owns software licensing. The other is a variable cost driven by demand, reviewed monthly, owned by whoever owns the service it powers. Merging them means neither one is actually being governed, and the first symptom of that is a surprise.

What to do about it

Three things, none of them difficult.

Split the line. Licences and consumption in separate budget lines, with separate owners, from the beginning. Retrofitting this after a surprising invoice is possible but the conversation is worse.

Ask for cost per interaction, not just total spend. Total spend rising is not information. Total spend rising while cost per interaction falls is a service getting more popular and more efficient. Total spend rising while cost per interaction also rises means something has changed in the design, and you want to know what.

Review licence assignment quarterly. Seats issued versus seats actually used, and reclaim the difference. It is the least interesting recommendation in this series and reliably the fastest saving.

How CloudNala can help

Where we tend to be useful is in producing the per-interaction number before a service launches rather than after — modelling what a realistic month looks like at expected volume, and being explicit about what happens at three and ten times that. It is not a difficult calculation. It is simply one that very few business cases contain, which is why so many of them are approved on a number nobody can defend six months later.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za