Platform Engineering24 August 20266 min read

Building Governed AI Agents — Part 14 of 17

Four AI engineering metrics teams should measure every week

Cost, quality and acceptance have to be read together and weekly. Any one of them alone will let you declare victory while the other two get worse.

#AI Engineering#Platform Engineering#FinOps#Engineering Leadership

Most organisations running AI-assisted engineering have either no regular measurement or a monthly cost figure. Monthly is too slow: a behaviour that started in the first week has four weeks of compounding before anyone sees it, and by then nobody remembers what changed.

Weekly is the right cadence, and four numbers are enough.

The weekly review, in five columnsCostQualityAcceptanceRiskActionsSpend this weekCost per acceptedunit of workFirst-passacceptance rate,by roleWhat passedreview, and whatcame backCache hit rateTop ten sessionsby costOne change,owned by onenamed personRead together, weekly. Any one of these on its own will mislead you.
The last column is what makes it a review rather than a report. A dashboard nobody has to act on gets read for about three weeks.

Cost per accepted unit of work

Total spend divided by the outputs that actually passed review.

This is the only cost metric that cannot be gamed by getting worse. Total spend falls if you use a cheaper model that produces work needing three attempts. Cost per turn falls if turns get less useful. Cost per accepted unit of work only falls if you are genuinely more efficient at producing acceptable output.

Defining the unit is the part that takes a conversation. A merged pull request, a completed ticket, an approved document, a signed-off compliance matrix: whatever it is, it needs to be something a person accepted, and it needs to stay constant so the series means something. It does not need to be perfect, it needs to be stable.

Cache hit rate

Watch it the way you watch an error rate: the level is uninteresting, the change is the signal.

A sudden drop means the shape of context construction changed, possibly because of a tooling update you did not initiate. That is worth ten minutes of investigation.

A high or rising value means nothing about efficiency, for reasons covered in why a 98 percent hit rate can hide an expensive workflow. Never put this number in front of leadership on its own. Paired with total volume it is informative; alone it is actively misleading.

First-pass acceptance rate, by role

What proportion of outputs are accepted without rework, split by the kind of work.

Both directions are informative, which is what makes this a good metric.

If it is very high, that is not unambiguously good news. It may mean expensive models are being used for work that does not need them, or that review has become a formality. A first-pass acceptance rate of ninety-five percent on complex architecture work is more likely to indicate a review problem than an exceptional agent.

If it is low, the instinct is to blame model strength. Inspect the briefs first. In our experience the majority of low acceptance rates trace back to under-specified requests rather than model capability, which is a workflow spec problem rather than a procurement one.

Splitting by role matters because a single blended figure hides everything. Simple code changes and architecture work have genuinely different acceptance profiles, and averaging them produces a number that describes neither.

Top ten sessions by cost

Not a summary of them. The sessions, read by a person, weekly.

This is the one that gets skipped, and it is the one that produces the insight. The other three tell you the state of things; this one tells you why. A team that reads its ten worst sessions every week for a month will have found and fixed most of its structural cost problems, without anyone writing a policy.

Ten minutes is enough. One person is enough. Rotating who does it is a good idea, because the point is partly to spread awareness of what expensive looks like.

Reading them in pairs

Each of these numbers has been used, somewhere, to declare success while something else deteriorated.

Metrics that only mean something in pairsTotal spendCache hit rateFirst-pass acceptanceAccepted units of workContext delta per turnWhich model was usedCost per unit that survived reviewWhether the discount hides growthWhether expensive models do easy work
Every number on the left has been used at some point to declare victory. None of them can do it alone.

This is why the review is a single conversation rather than four separate dashboards. Total spend fell and acceptance fell further: that is not a saving. Cache hit rate rose while context per turn also rose: the discount is masking growth. First-pass acceptance is excellent and the most expensive model is doing routine work: you are buying certainty you did not need.

None of those states are visible from one number.

The fifth column

The review needs an output or it stops being a review. One change, owned by one named person, before the next one.

Not a list of observations. Not a set of recommendations for the team to consider. One change, with a name against it, that can be checked next week. A dashboard nobody has to act on gets read attentively for about three weeks and then becomes furniture.

This is the same discipline that makes any operational review work, and AI spend is not special enough to be exempt from it. Our note on cloud cost governance makes a structurally identical argument about a different bill.

What good looks like

Four numbers, weekly, in one conversation rather than four dashboards.

Cost expressed per accepted unit of work, with the unit defined and stable.

Acceptance split by role, never blended.

The top ten sessions read rather than summarised.

One change, one owner, one week, checked at the next review.

Practical checklist

  • Define the accepted unit of work and keep it stable
  • Report cost per accepted unit, not total spend
  • Split first-pass acceptance by role
  • Read the ten most expensive sessions, do not just chart them
  • Pair every metric with the one that stops it lying
  • End every review with one change and one name
  • Keep it to thirty minutes, weekly

How CloudNala can help

The hard part of this is rarely the data, it is agreeing what an accepted unit of work is and then holding the cadence. We help teams set the four metrics up, pair them so no single number can be used to declare victory, and establish the weekly rhythm with an owner. The first three or four reviews usually pay for the whole exercise.

The metric set here follows Andrew Baker's August 2026 recommendations for instrumenting coding agents.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za