CloudNala Builds24 August 20268 min read

Building Governed AI Agents — Part 17 of 17

The CloudNala production agent playbook

Fourteen steps from one workflow to a system a business can depend on, in the order they pay off. Only the first one is needed to start.

#AI Agents#CloudNala Builds#AI Governance#Solution Architecture

Everything in this series describes one piece of a system. This article puts them in order.

The order matters more than the list. Most agent projects that stall did not skip a step, they did them in a sequence where each one had to be redone: evals written before the workflow was stable, memory added before anyone knew which facts mattered, expansion decided before quality could be measured.

The reference shape of a governed agentMEMORY LAYERApproved facts, decisions and next actions, written through a gate, with expiryBusinesseventWorkflowspecReasoningand toolsHumanreviewOutputGUARDRAILSPermissionsApproval gatesSource limitsCost ceilingsEscalation rulesOBSERVABILITYEvals · Traces · Per-turn cost and context · Alarms and ceilings · Weekly review
Memory above, guardrails beside, telemetry beneath, and a human inside the flow rather than bolted onto the end of it. Everything in this series is one of these four blocks.

The four blocks

The shape above is the reference we come back to on most engagements, and it is deliberately unremarkable.

The flow runs left to right. A business event arrives. A workflow spec says what to do about it. Reasoning and tools do the work. A human reviews. Something is produced. The human sits inside the flow rather than bolted onto the end, which is the difference between review as a control and review as a formality.

Memory sits above it, feeding every step. Approved facts, decisions and next actions, written through a gate and with an expiry. Not conversation history.

Guardrails sit beside it, applying throughout: permissions, approval gates, source limits, cost ceilings, escalation rules. They are drawn as a rail rather than a step because they are not a stage the work passes through, they are a boundary the work stays inside.

Telemetry sits underneath, capturing from every stage: evals, traces, per-turn cost and context, alarms, the weekly review. It is the foundation in the diagram because it is what everything else is inspected through.

Cost observability belongs in that bottom layer rather than in a separate finance system, which is the one structural argument this series makes that is not obvious at the start.

The fourteen steps

Fourteen steps, in the order they pay offFrameBuildInstrumentImprove1 · Choose one vertical2 · Define the outcome3 · Meet users in thetools they already use4 · Write the spec5 · Curate the examples6 · Add smart defaults7 · Define memory rules8 · Add the human loop9 · Build the eval set10 · Instrument per-turncost and context11 · Set alarms andcost ceilings12 · Read top sessionsweekly13 · Improve spec,examples, guardrails14 · Expand only aftermeasured stabilityExpansion is the reward for stability, not the method of achieving it
Only step one is needed to start. The rest describe what has to become true before a pilot is allowed to become a dependency.

Frame

1. Choose one vertical workflow. One, and narrower than feels ambitious. Why vertical first.

2. Define the business outcome. Not "assist with tenders" but "reduce the time to a bid or no-bid decision from three days to half a day, without missing a disqualifying condition". The second clause is the one that makes it a real target.

3. Meet users in their current tools. Version one changes the worst step, not the interface. Where the work already lives.

Build

4. Write the natural-language workflow spec. Numbered steps, each with an input, an output and a "cannot answer" branch. Specs, not vibes.

5. Create curated examples. Five categories: accepted, corrected, edge cases, rejected, anti-patterns. Examples are the fuel.

6. Add smart defaults. Business actions on the front page, settings behind an expert layer. Smart defaults.

7. Define memory rules. Remember, ignore, expire, correct, audit. Start with no memory and add only what breaks without it. Memory as a designed layer.

8. Add the human correction loop. With reason codes that route somewhere, not a thumbs-down. Human review as the improvement mechanism.

Instrument

9. Build evals. Twenty real cases with known answers, wired to a release gate. Continuous evals.

10. Instrument per-turn cost and context. Turn, model, input size, context delta, cache, estimated cost. Instrument every turn.

11. Set alarms and cost ceilings. A context-delta alarm catches the specific action that causes most of the damage, on the day it happens.

Improve

12. Review top sessions weekly. Read them, do not summarise them. Four weekly metrics.

13. Improve the workflow, examples and guardrails. One change, one owner, checked at the next review.

14. Expand only after measurable stability. Expansion is the reward for stability, not the method of achieving it.

Where teams actually go wrong

Three patterns, seen often enough to be worth naming.

Starting at step 14. The platform is designed before the first workflow works. Every subsequent decision is made in the abstract, and the first real use case turns out to need something the platform made hard.

Doing 9 through 11 last. Evals and telemetry are treated as hardening work for after the pilot. By then there is no baseline, no known-good set, and no way to tell whether the last three changes helped. This is the most common single cause of an agent project that cannot get past pilot.

Skipping 8 because it looks like a bottleneck. The human loop gets designed as a temporary approval step, so it collects no reasons, so it generates no data, so nothing improves, so it genuinely does become a bottleneck. The prophecy fulfils itself in about four months.

The honest scope of this

Fourteen steps is a lot to put in front of a team that wants to try something. It is worth being clear that this is the shape of a system a business will depend on, not the shape of a first experiment.

If you are exploring, do step 1 and step 4, and build the thing. That is a legitimate and useful piece of work, and it will tell you whether there is anything here.

The steps in this list are what has to become true before that experiment is allowed to become a dependency. The distinction between those two states is the one most organisations manage badly: something built as an experiment becomes load-bearing without anyone deciding that it should, and the governance never catches up. We have written about that transition in the trust checklist.

Practical checklist

  • Pick one workflow and write the outcome down with its quality constraint
  • Write the spec before writing any prompt
  • Build the example library and the review loop in the same pass
  • Get evals and per-turn telemetry in before the pilot, not after
  • Set a context-delta alarm on day one
  • Hold a thirty-minute weekly review with one owner and one change
  • Decide explicitly when the experiment becomes a dependency, and say so

How CloudNala can help

We work through this playbook with clients in roughly this order, and the early steps are usually facilitation rather than engineering: getting the workflow written down properly, getting the examples out of people's folders, and agreeing what the outcome actually is. The engineering follows more easily once those exist. Where we are most often useful is at steps 9 to 11, because they are the ones teams intend to do later and rarely get to.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za