The prototype went down well. It always does. Someone had built a citizen-facing assistant on a rapid-prototyping platform, and in the room it looked finished: the flows worked, the language was right, it answered questions about services in a way that felt genuinely useful. Everyone agreed it should be built.
Then we tried to put it into the client's own cloud tenant, and discovered the demo had told us almost nothing about that.
This is the first article in a series about the distance between those two moments. It is drawn from a real delivery for a South African public-sector organisation, anonymised throughout. The methodology is the point; the client is not.
What a prototype is actually evidence of
A prototype is a genuine, valuable artefact. It settles arguments that are otherwise very expensive to settle: whether the flow makes sense, whether the tone is right, whether people can find what they came for, what the content needs to look like. Those are real questions and a demo answers them well.
What it does not answer is anything on the right-hand side of that diagram. It does not tell you whether the thing can be deployed into a tenant governed by someone else's security policy. It does not tell you how identity works when the deploying team are external guests. It does not tell you where the data is legally allowed to live. And it says nothing at all about who runs it once the delivery team has gone.
The failure mode is not that people think the prototype is the product. Most people know it isn't. The failure mode is subtler: the prototype makes the remaining work look like a port. It looks like you take the thing that exists and move it. So the remaining work gets estimated as a port, and the estimate is wrong by a lot — not because the engineering is hard, but because almost none of the remaining work is engineering.
Harvest the content, regenerate the architecture
The instinct with a working prototype is to inherit it. Take the codebase, adapt it, keep going. On a rapid-prototyping platform this instinct is close to always wrong, and it is worth being precise about why.
The valuable part of a prototype is the part that cost human time and judgement. Copy that has been through a review cycle. Translations. The shape of the data. A theme somebody agonised over. Flows that real users have already reacted to. All of that is expensive to recreate and should be lifted wholesale.
The part that looks most like the product — the runtime, the storage, the auth, the hosting — is the part the prototyping platform gave you for free, and it is the part that will not survive contact with a regulated tenant. It was never yours. It is not designed for the deployment model, the identity model, or the residency constraints you are about to meet. Carrying it forward means spending weeks adapting something you would not have chosen, to reach a place you could have started from.
So: treat the prototype as a specification input and a visual reference. Regenerate the architecture clean against the target platform. This sounds wasteful and is the opposite.
Say "deployed" only when it is deployed
The second lesson from this phase is a reporting discipline rather than a technical one, and it caused us more trouble than any code did.
There is a large and easily-blurred difference between these two statements:
- The artefacts are ready.
- The thing is running in the client's tenant.
On a milestone-based programme, conflating those two is a delivery risk and an evidence risk at the same time. It is a delivery risk because "ready" work that cannot be deployed can sit for a fortnight while an access request moves through someone's queue, and nobody notices the clock running because the status said green. It is an evidence risk because the artefact that proves a milestone is usually the deployed thing — a URL, a resource view, a pipeline run — and you cannot produce it retrospectively for a thing that was never deployed.
We ended up separating the two explicitly in every status update: what is built, and what is live. It reads as pedantic for about a week and then starts catching things.
This is where projects actually die
It would be easy to read the above as a story about one awkward deployment. The numbers suggest otherwise.
Writing in ITWeb in August 2026, Eugene Perumal of Valutivity (Deploying AI agents safely) cites research putting agentic AI project failure at around 40%, attributed not to model quality but to missing foundations — the governance, access and accountability work that the demo never touches. The same piece reports that 95% of executives say their organisations have already experienced negative consequences from enterprise AI use, with direct financial loss the most common outcome, and that 65% of enterprise leaders cite agentic system complexity as their top deployment barrier.
His conclusion is the one worth carrying into the rest of this series: "governance maturity is the single strongest predictor of AI readiness." Not model choice. Not the quality of the prototype. The maturity of the boring machinery around it.
That reframes the gap we are describing. The distance between the demo and the delivery is not an inconvenient tail of admin at the end of an AI project. On the evidence, it is the AI project — and treating it as overhead is roughly how two in five of these programmes end.
What this means for how you budget
The practical consequence is that an AI programme has two distinct cost centres and most estimates only contain one of them.
The first is building the thing — and with a good prototype plus modern tooling, this is now the cheaper and more predictable half. That is a genuine change from five years ago and it is worth saying plainly.
The second is everything required to run it inside someone else's governed environment: access, identity, least-privilege, data residency, evidence, operability, content. This half has not got cheaper. If anything it has got more involved, because regulated organisations have tightened their cloud governance considerably over the same period that building got easier.
Budget for the second half explicitly. Give it its own line, its own owner, and its own dates. The rest of this series is a walk through what is actually in it: the access chain nobody demos, least-privilege as code, where your AI is allowed to run, evidence-first delivery, and why content is usually the real critical path.
What we would do differently
Three things, all cheap, all things we now do by default:
Write the access map before writing infrastructure. Every approval you will need, listed, before it blocks you. That article is next in the series.
Deploy a walking skeleton in week one. One thin slice, all the way through, into the real tenant. It proves the entire right-hand column of that first diagram while there is still time to react.
Split "built" from "deployed" in reporting from day one. Not as a caveat at the bottom of a slide — as two separate columns.
How CloudNala can help
Most of the organisations we meet in this position have already done the expensive, creative part. They have a prototype people liked and a genuine appetite to ship it. What they are missing is a realistic read of the second half: what the tenant will demand, which approvals will bite, where the data is allowed to sit, and what evidence the programme will be asked for later. That read is usually a short piece of work, and it is considerably cheaper before the delivery estimate is written than after.
Work with CloudNala
CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.
Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.
Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za