Platform Engineering24 August 20266 min read

Building Governed AI Agents — Part 15 of 17

AI cost observability needs three views: developer, team and finance

Three audiences ask three different questions about the same spend. Building one dashboard and expecting it to serve all of them is how the numbers lose credibility.

#AI Engineering#Platform Engineering#FinOps#Observability

There is a recurring moment in AI cost conversations where an engineer says the workflow costs about R400 a run and finance says the invoice does not support that. Both are right. They are measuring different things for different reasons, and neither number is wrong.

The mistake is upstream of the disagreement: treating AI cost observability as one problem with one dashboard.

Three views of the same spendLAYER 1Local telemetryLAYER 2Provider billingLAYER 3Organisation analyticsWhy did this session cost so much?Read by: the developer, architect or agent builderWhat was actually billed, to the cent?Read by: finance, procurement, the platform ownerWhich teams or workflows are trending unusually?Read by: engineering leadership, CIO, governance
No single dashboard answers all three questions, and the common failure is building the middle one and assuming it covers the other two.

Local telemetry answers "why did this cost so much"

Audience: the developer, the architect, the person who built the agent.

This is estimated, per-turn, immediate, and its value is diagnostic. It tells you that context stepped up at turn 43 and that the session ran for four days. It does not need to be accurate to the cent, and insisting that it should be is the most common way this layer never gets built.

The property that matters here is proximity. The person who caused the cost sees it close enough in time to connect it to what they did. That feedback loop is what changes behaviour, and no amount of monthly reporting substitutes for it.

Provider billing answers "what were we actually charged"

Audience: finance, procurement, the platform owner.

This is authoritative, delayed, and aggregated. It is the number that goes in the accounts, and it is the only one that settles a dispute.

It is also nearly useless for diagnosis. By the time it arrives, the sessions that caused it are weeks old and nobody remembers them. Asking finance to explain a spike is asking the wrong layer.

The useful discipline is periodic reconciliation: check that the estimates in layer one track the invoice in layer two, within a tolerance everyone has agreed. They will not match exactly and they do not need to. What matters is that the gap is stable and understood, because that is what allows engineering estimates to be trusted in a conversation where the invoice is not yet available.

Organisation analytics answers "what is trending unusually"

Audience: engineering leadership, CIO, governance.

This is the view that is most often missing, because it only becomes necessary at a certain scale and by then the other two are entrenched.

It answers questions neither of the others can: which teams are growing usage fastest, which workflows have changed shape, whether adoption is broad or concentrated in a few people, whether the value being reported justifies the trend. It is comparative rather than absolute, and it is read for direction rather than for detail.

What breaks when one is missing

What breaks when one view is missingNo local telemetryNo billing reconciliationNo organisation viewYou know the bill. You cannotexplain which turn caused it,so every fix is a guess andno fix can be proven.Estimates drift from theinvoice. Finance quietly stopstrusting any number thatcomes from engineering.One team triples its usageand nobody notices untilthe quarter closes and thevariance needs explaining.Three questions, three audiences, one shared source of truth underneath
Each of these is a real failure mode rather than a hypothetical one, and each is invisible from inside the other two views.

The middle failure is the one that does lasting damage. Once finance has been given an engineering number that turned out to be materially wrong, every subsequent engineering number is discounted. Rebuilding that credibility takes considerably longer than setting up the reconciliation would have.

The third failure is the one that produces surprises at quarter end, which is the worst possible time to discover a trend that started three months earlier.

One source, three views

The important architectural point is that these are three views, not three systems.

The per-turn record captured at the source contains what all three need: session, user, workflow, model, tokens, cost estimate, timestamp. The developer view filters it to one session. The organisation view aggregates it by team and workflow over time. The finance view reconciles the aggregate against the invoice.

Building three separate collection mechanisms produces three numbers that disagree, and then a project to work out why. Capture once, at the source, and derive the rest. This mirrors the argument we make about observability for AI workflows generally: one trace, several audiences.

Cost estimates, billing truth, governance direction

A short way to remember which is which, and it is worth saying to stakeholders explicitly at the start:

Cost estimates explain behaviour. Billing confirms money. Organisation analytics guides governance.

Nobody should be asked to make a decision from the wrong one. Do not ask finance why a session was expensive. Do not ask a developer to confirm the quarterly figure. Do not try to spot a three-month trend from one session's telemetry.

What good looks like

One capture at the source, three derived views.

Estimates accepted as estimates, with an agreed tolerance against billing.

A periodic reconciliation that somebody owns, so the gap stays understood.

The organisation view built before it is urgently needed, not after the first quarter-end surprise.

Each audience looking at the view built for their question.

Practical checklist

  • Capture per-turn data once, with session, user, workflow and model attached
  • Derive the developer, team and finance views from that single record
  • Agree a tolerance between estimates and billing, and check it monthly
  • Give the reconciliation an owner
  • Build the organisation view before scale makes it urgent
  • Tell each audience which view answers their question, and which does not

How CloudNala can help

Most teams have some version of layer one and all of layer two, and nothing joining them. We help design the capture so all three views derive from a single record, set up the reconciliation that keeps engineering estimates credible with finance, and build the organisation view before it is needed rather than after a surprise. The technical work is modest; the value is that the numbers stop being disputed.

The three-layer distinction here follows Andrew Baker's August 2026 article on instrumenting coding agents.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za