There is a particular kind of AI product demo that goes well in the room and badly in the building. It has a text box, a model selector, a few sliders for retrieval settings, and an invitation to describe what you want. Everyone nods. It looks flexible, and flexibility sounds like a feature.
Then it ships to twenty people in a procurement team, and within a fortnight three of them are using it properly, four have given up, and the rest are pasting in a prompt that somebody wrote in week one and forwarded around on email. Nobody is quite sure whether the answers are good. The prompt has been edited by four people and nobody can say what the edits were for.
The problem is not that the users are unsophisticated. The problem is that the product asked them to supply the one thing it should have brought with it: the workflow.
Your users are worse prompt engineers than you, and that is fine
This sounds harsh written down. It is simply true, and it is true for a reason that has nothing to do with capability.
You have spent months with the model. You know that it responds better when you name the output format, that it fabricates less when you tell it what to do with uncertainty, that a request phrased one way gets a summary and phrased another way gets an analysis. That is accumulated craft, and there is no reason a compliance officer should have any of it.
They know something you do not, which is what a defensible compliance answer looks like at their organisation. The design question is how to get those two kinds of knowledge into the same place. Asking the compliance officer to become a prompt engineer is one option. It is the slow one.
What a smart default actually is
A default is not a shortcut, and it is not a preset. It is a decision the product has made on the user's behalf, because the product is in a better position to make it.
For a procurement user, the difference is between being asked to write this:
Act as a tender compliance analyst. Extract the mandatory requirements from the attached document, identify any disqualifying conditions, map each requirement to supporting documents in our library, and flag anything you are unsure about.
and being offered this:
Check compliance
Behind that button sits the same instruction, written once by someone who knows how the model behaves, reviewed by someone who knows what compliance means here, versioned, and testable. When it improves, everyone gets the improvement. When it is wrong, there is one place to fix it.
For TenderCity, that turns into a short set of actions rather than a text box: check tender eligibility, create a compliance matrix, summarise mandatory requirements, draft clarification questions, prepare a bid or no-bid note. Each one is a workflow that a bid team already recognises. None of them require anyone to describe what they want in prose.
The part people get wrong
The argument here is not that customisation is bad. It is that customisation belongs somewhere specific, and the front page is not it.
Three layers, three different populations. The business user gets actions. The power user gets a small number of adjustments that are visible but deliberately not on the default path: change the source set, adjust the output format, re-run one step. The product team gets everything else, and changes it with evidence rather than intuition.
That third layer is where the discipline lives. Model choice, retrieval parameters, prompt structure and tool wiring are genuine engineering decisions with genuine consequences, and they should be changed the way any other production configuration is changed: deliberately, by someone accountable, and tested against a known set of cases before it reaches anyone.
The failure mode when this collapses is quiet. Every user's results drift apart, because every user has a slightly different prompt. Nobody can compare two outputs, because they were produced by two different systems. And the moment somebody asks "is this thing actually any good?", there is no answer, because there is no thing — there are twenty of them.
Some black-box stability is a feature
There is a reflex in engineering culture that says exposing more control is more honest. In an AI product, some opacity is what makes the output comparable.
If two people in the same team ask the same question and get materially different answers because one of them phrased it better, the system has no reliability to speak of. You cannot build a review process on it, you cannot write an eval for it, and you cannot tell a client what it does. Fixing the workflow in place is what makes it possible to say: this is what the system does, here is how we test it, here is what it does when it is unsure.
That stability is also what lets you improve. A workflow that lives in one place can be measured, changed, and measured again. Twenty prompts in twenty inboxes cannot.
What good looks like
A short version, for a team about to build this.
Start from the business action, not the capability. "Check compliance" is an action. "Ask the document anything" is a capability, and capabilities are what you offer when you have not decided what the product is for.
Make the agent ask for what it needs. If a workflow requires the evaluation criteria and they are not in the uploaded pack, the right behaviour is to ask, not to guess and carry on confidently.
Keep the output shape fixed. A compliance matrix should come back looking like a compliance matrix every time, because that is what makes it reviewable by someone who did not run it.
Put the settings behind a door, and put someone's name on the door. Someone owns the prompt, the source set and the model choice. When the output degrades, there is a person to ask and a change history to read.
Then write the workflow down properly, which is the subject of the article on workflow specs, and give it a way to be corrected, which is the human review loop.
Practical checklist
- List the five business actions your users actually need, in their words
- Write one workflow spec per action, owned by a named person
- Decide what the agent does when an input is missing, before it ships
- Fix the output format so two runs can be compared
- Move model, retrieval and prompt settings behind an expert layer
- Test every default against a known set of cases before changing it
How CloudNala can help
We usually start these engagements by watching people work rather than by looking at the model. The five actions worth building are almost always already visible in what a team does manually every week, and they are rarely the ones described in the original brief. From there the work is ordinary product design: decide the defaults, write them down, put them behind a small number of buttons, and give one person ownership of what sits behind each.
Work with CloudNala
CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.
Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.
Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za