Cloud Strategy9 September 20267 min read

Why Do We Keep Paying for AI? — Part 7 of 8

Governing the AI bill: FinOps for AI

The AI bill is not a runaway train — but only if someone is driving it. The controls are unremarkable; what is usually missing is an owner, six honest numbers and the discipline to scale spend with value.

#FinOps#Cloud Cost#AI Governance#Leadership

Everything in this series so far explains why the AI bill behaves the way it does. This article is about keeping it under control, and the short version is that the controls are not the hard part.

The controls are well understood, unremarkable, and mostly borrowed from cloud financial management, which most organisations have already had a go at. What is usually missing is not a technique. It is an owner.

Three views of the same bill, and the number that joins themFINANCESees the invoiceCannot tell whichworkload caused itIT AND ENGINEERINGSees the usageNot accountable forwhat it is worthTHE SERVICE OWNERSees the valueUsually has no ideawhat it costs to runTHE ONE METRIC ALL THREE CAN HOLDCost per outcome — per resolved query, per served customer, per processed claimWith an owner, the bill is a monthly management decisionWithout one, it is a quarterly discovery
Each of these three is holding a third of the picture and can defend their own third completely. The failure is structural, not personal: no single one of them can answer whether the spend was worth it.

Nobody in the room can answer the question

Here is the structural problem, and it is worth being precise about because it is not anybody's fault.

Finance receives the invoice. They can see the total, and they can see it moving. They cannot tell which service caused which portion of it, or whether the movement was good news.

IT and engineering can see the usage in detail. They know exactly which workload consumed what. They are not, however, accountable for whether that workload was worth running.

The service owner — the person responsible for the customer journey the AI sits inside — knows precisely what it is worth. They usually have no visibility of what it costs, and often no idea that they could.

Three people, three thirds of the picture, each able to defend their own third completely. And no one of them can answer the only question that matters: was that spend worth it?

This is why AI cost governance fails in organisations that are perfectly competent at cost governance generally. The failure is not analytical. It is that the analysis has no home.

The controls, briefly

With an owner in place, the practical controls are short and mostly obvious. They are listed here in the order they should be applied, because applying them in the wrong order is how organisations end up retrofitting limits onto services people already depend on.

Right-size at design time. Choose the model the job needs, not the largest one available. Trim what gets attached to each request. Cache answers to questions that repeat — and in any real service, a large share of questions repeat. All of this is dramatically cheaper to do at design time than to retrofit, because retrofitting means changing something users have already come to rely on.

Launch small and scale on evidence. Start on demand. Start with a subset of users or a single journey. Let the usage pattern reveal itself before committing to anything with a monthly floor.

Instrument before you launch, not after. Cost attributed per service, per journey, and ideally per interaction — visible from day one. A service that has been running for six months without cost attribution cannot be optimised, only guessed at.

Set budgets and alerts. A threshold with an alert attached, agreed before launch, so that an unusual month produces a notification rather than a discovery. This is trivial to configure and startlingly often absent.

Review commitments on a schedule. Any reserved capacity gets a monthly utilisation check for the first year. A commitment nobody is watching is a subscription nobody cancels.

Measure cost per outcome, not just total cost. The number that makes every other number interpretable.

The six numbers

If a leader takes one practical thing from this series, let it be this list. These are the numbers to require on a monthly dashboard — and requiring them consistently will surface every failure mode in this series early enough to act on.

Six numbers to ask for every monthSPEND, SPLIT BY KINDLicences on one line, consumptionon another. Never one merged total.COST PER OUTCOMEPer resolved query, per servedcustomer. The only honest unit.VOLUME TRENDIs the bill rising because use isrising, or because something broke?RESERVED CAPACITY IN USEWhat fraction of the lane you arepaying for is actually being drivenLICENCES ISSUED VS USEDThe quietest waste in the building,and the easiest to reclaimBUDGET, ACTUAL AND AN ALERTSet before launch, not after thefirst surprising invoiceSpend should track value. These are the numbers that show whether it does,and they only work if somebody is required to present them.
Six is deliberate. A leader who asks for these consistently will catch every failure mode in this series early enough to do something about it, and will not need to understand a single thing about how the models work.

A note on each of the two that get argued about.

Spend split by kind matters because merging licence and consumption costs into one AI total destroys the information in both. They move for different reasons and are managed by different means. One merged number tells you nothing except that it went up.

Cost per outcome is the one people resist, because agreeing the denominator requires a conversation about what the service is actually for. That conversation is the point. An organisation that cannot say what one unit of output from its AI service is worth has not finished designing the service.

Who owns it

Not IT alone. Not finance in the dark. Both, with the service owner in the room.

In practice the arrangement that works is unremarkable: a standing monthly review, half an hour, with a named chair. Finance brings the spend. Engineering brings the usage and the technical explanation for any movement. The service owner brings the outcomes. The dashboard is the same six numbers every month, and somebody is answerable for each of them.

What makes it work is not the meeting. It is that the movement in the bill has to be explained by someone who understands both halves. "Consumption rose eighteen percent" is not an explanation. "Consumption rose eighteen percent because resolved queries rose twenty-two percent, so cost per resolved query fell" is an explanation, and it is also good news that would otherwise have looked like a problem.

What good and bad look like

It is worth being concrete about what the review is looking for, because "the bill went up" is not by itself a finding.

Healthy. Total spend rising while cost per outcome falls or holds. Reserved capacity utilisation high. Licence assignment close to licence usage. Volume growth traceable to something you did on purpose.

Worth investigating. Cost per outcome rising — something changed in the design, or the service is escalating more than it resolves. Reserved capacity utilisation below half — you are paying for a lane you are not driving on. A large gap between licences issued and licences used — the quietest waste in the building and the fastest saving available.

Actually alarming. Spend rising with no corresponding rise in usage at all. That is usually a defect: something retrying in a loop, a runaway job, or a misconfiguration. It is rare, and it is the reason the alert threshold exists.

Scale spend with value, not ahead of it

The single sentence that summarises this article, and arguably the series: let the spending follow the evidence.

Almost every expensive AI mistake in the last two years has been an organisation committing ahead of proof. Buying licences for everyone before knowing whether anyone would use them. Reserving capacity before the traffic existed. Building for a scale that had not arrived. In every case the money went out before the evidence came in, and in every case the evidence, when it arrived, would have suggested something smaller.

The discipline is not to spend less on AI. Organisations that under-invest here will find that out too, more slowly and more painfully. The discipline is to make each increment of spend conditional on the previous increment having produced something.

That is not a technology practice. It is ordinary management applied to an unfamiliar cost shape.

How CloudNala can help

We set up the review rather than run it forever: the six numbers, wired to real data, with an owner named for each and a threshold that alerts before anyone is surprised. The pattern we most often correct is a service six months live with no cost attribution at all — where the only available lever is a blunt one, because nobody can see which part of the bill is which.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Request a Cloud Review or write to us at consult@cloudnala.co.za