AI Strategy19 August 20269 min read

From Chatbot to Agent — Part 10 of 15

From local prototype to production-ready AI agent

The demo took a weekend. Getting it into production has taken seven months and it is still not there. That gap is not a technology problem — it is everything a prototype is allowed to leave out.

#AI Strategy#AI Governance#Platform Engineering#Delivery

There is a conversation that happens in South African organisations roughly every second week now. Someone built something impressive with AI in a few days. Everyone who saw it agreed it was impressive. That was five months ago and it is still not in production, and nobody can quite explain why.

The explanation is usually the same, and it is not about the technology. A prototype is permitted to assume a great deal: that the data is clean, that the user is the person who built it, that failures can be fixed by hand, that nobody is accountable for the output, that cost does not matter, and that when it breaks someone will notice.

Production permits none of those.

From local demo to monitored service1Localdemo2Internalprototype3Controlledpilot4Human-approvedworkflow5Productiondeployment6Monitoredservice7ContinuousimprovementGovernance, ownership and cost control increase at every step — not only at the end
Most AI projects that stall are stuck between steps 2 and 3 — not because the technology failed, but because nobody defined who owns it, who approves its output, or how anyone would know it was wrong.

Where projects actually stall

Almost all of them stall between the internal prototype and the controlled pilot. That transition is where the questions stop being technical.

The gate AI projects actually stall atAn impressiveinternal prototypeTHE ACTUAL GATEWho owns it, and is that funded?Which data sources are approved?Who approves the output?How do we know it works?What stops the cost growing?What happens when it breaks?A controlledpilotThese are the ordinary conditions for putting anything into production. AI prototypes just make them easy to skip.
Not one of these is a technology question, which is why a stalled project rarely gets unstuck by better engineering. Narrow the workflow and every answer gets shorter.

Who owns this? Not who built it — who is accountable for it working, funded to maintain it, and answerable when it is wrong. An AI system without a named business owner will not clear a risk review, and it should not.

Which data sources are approved? The prototype used a folder someone had access to. Production needs a defined, owned set of sources with permissions that reflect the actual entitlements of the people using it.

Who approves the output? Where the output carries consequence, someone signs it. That person needs to see enough to make the decision meaningfully, and their approval needs to be recorded.

How do we know it works? Not impressions — the evaluation set covered in the article on evaluation. This is frequently the single blocking item, and it is one of the cheaper ones to resolve.

What does it cost, and what stops it costing more? A ceiling enforced in code, and an alert on cost per case.

What happens when it breaks? Who is called, what the fallback is, and how the work gets done while it is down. If the answer is "we go back to doing it manually", that is a perfectly good answer — but it needs to be written down and the manual path needs to still exist.

None of these are AI questions. They are the ordinary requirements for putting anything into production in a regulated, audited organisation. AI prototypes just make it unusually easy to skip them, because the thing appears to work so early.

The production checklist

Before an AI workflow goes live, we would want each of these answered in writing:

  • A named business owner, funded to maintain it
  • A defined workflow with the decision points marked
  • Approved data sources with owners and review dates
  • Access control inherited from source systems, applied at retrieval
  • Human approval gates on every irreversible action
  • An evaluation suite with a recorded baseline
  • Traces captured on every run, with redaction decided
  • A cost ceiling per task and an alert on cost per case
  • A rollback plan and a working manual fallback
  • A named support owner and an incident path
  • Documentation a new team member can follow
  • Security review completed
  • Privacy and POPIA review completed
  • An agreed measure of business value, with a baseline taken before go-live

That last item is worth insisting on. If nobody measured how long the process took before the AI arrived, nobody will be able to prove it improved, and the system will be evaluated on anecdote. Take the baseline first — it costs an afternoon and it is unrecoverable afterwards.

Doing this without stopping everything

A checklist of fourteen items can read as a reason not to start. It is the opposite: it is a reason to start smaller.

Every item on that list is easier to satisfy for a narrow workflow than a broad one. A system that classifies incoming documents and populates a checklist has a small blast radius, a clear owner, an obvious evaluation set and an easy fallback. A system that "handles procurement" has none of those, which is why it will spend a year in review.

The organisations that get AI into production fastest are not the ones with the highest risk tolerance. They are the ones that scoped the first workflow narrowly enough that the governance questions had short answers.

The staged path

Local demo proves the idea is possible. Its only job is to answer "could this work at all". Keep it to days.

Internal prototype puts it in front of five people who do the work. Its job is to find out whether the output is actually useful, which is a different question and frequently answered no. Better to learn that here.

Controlled pilot runs it against real cases, in parallel with the existing process, with everything logged. Its job is to produce evidence: how often is it right, how often does a human change the output, what does a case cost. This stage is where the evaluation set and the traces earn their existence.

Human-approved workflow is the first version that is genuinely in the business. The agent prepares, a person decides. For many processes this is the end state, not a waypoint.

Production deployment means it is a system: supported, monitored, owned, with an incident path.

Monitored service means someone looks at the numbers regularly and acts on them.

Continuous improvement means new cases enter the evaluation set, the prompt and retrieval evolve against it, and model changes are tested rather than absorbed.

Most organisations should expect to spend real time between stages three and four, and should stop treating that as a delay. That is where the system becomes trustworthy.

A worked progression

A document intelligence capability might reasonably grow like this: start with question-answering over a document set, which proves the retrieval works. Add automatic summarisation on intake, which saves the first real hours. Add extraction into a structured checklist, which is the first time the output feeds another process. Add a recommendation with reasoning attached, still decided by a person. Then, once there is a year of traces showing where it is reliable and where it is not, consider which of the low-consequence steps can run without review.

Each stage is independently useful. If funding stops after stage two, the organisation still has something. That property — every stage shipping value on its own — is what keeps AI programmes alive through a budget cycle.

How CloudNala can help

We take AI prototypes through the stages that actually block them — establishing ownership, approving data sources, designing the approval gates, building the evaluation baseline, instrumenting cost and traces, and writing the rollback and support model. Often the technical work is nearly done and what is missing is the production wrapper, which is a matter of weeks rather than months once someone is driving it.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za