The horizontal pitch is always more attractive. A general business assistant addresses everyone, which makes the market look enormous and the roadmap look short. A tender compliance assistant for South African ICT tenders addresses a much smaller room, and sounds like a smaller ambition.
In practice the narrow one is more useful within weeks, and the broad one spends a year being nearly helpful at a large number of things.
The reason is not that general models are weak. It is that a general agent has to infer the shape of the problem on every single run, and inference is where the errors live.
What "vertical" actually means here
It does not mean picking an industry. It means picking a workflow that has a known shape, and building for that shape until it is boring.
An ICT tender in South Africa arrives as a recognisable object. There is a bid pack with a predictable set of documents. There are mandatory returnable schedules. There is a B-BBEE component with a scoring mechanism. There are compulsory briefing sessions that disqualify you if missed. There is a functionality threshold before price is even opened. There is a set of traps that catch inexperienced bidders every time.
None of that is in a general model's understanding of "a tender". All of it can be given to a narrow one, once, by people who know it.
That is the trade. A vertical agent gets told the things a general agent has to guess.
High-volume, domain-specific examples are what stop the silly mistakes
There is a stage every agent project goes through where the output is impressive in the demo and unreliable in the wild. The usual instinct is to improve the prompt. The thing that actually fixes it is volume of domain examples.
When a compliance agent has seen two hundred real bid packs, with the correct extraction marked up alongside them, it stops making the category of error that makes reviewers lose faith: reading a delivery date as a submission deadline, treating an optional annexure as mandatory, missing a disqualifying clause because it was phrased unusually in this one department's template.
Those errors are not fixed by better reasoning. They are fixed by having seen enough of the domain to know what the fields mean here. That is why the examples matter more than the model, which is the subject of the next article in this series.
Getting two hundred marked-up examples of tender packs is achievable. Getting two hundred marked-up examples of business documents is not a project, it is a category error.
Three vertical starting points
TenderCity: start with ICT tenders, not all tenders. Construction tenders, medical supply tenders and ICT tenders share a legal framework and almost nothing else in their evaluation logic. Doing one properly produces a system a bid team will rely on. Doing all three at once produces a system that is right often enough to be dangerous.
Lead triage: start with technology consulting opportunities. The signals that make a consulting opportunity worth pursuing are specific and learnable: the shape of the buyer, whether there is budget or only interest, whether the scope has an architecture decision inside it. A general lead scorer has to be told all of that in every prompt. A narrow one already knows.
Proposal and architecture support: start with one proposal type. AI and RFP architecture packs have a structure, a set of reusable patterns and a known set of weak claims that reviewers always push back on. That is a tractable target. "Automate proposals" is not.
Expanding without losing what you built
The staircase in the first diagram is not decoration. Each step is a precondition for the one above it, and skipping one is how horizontal expansion goes wrong.
You need repeatable examples before you have a stable workflow, because the workflow is really a summary of what the examples taught you. You need a stable workflow before quality can be measured, because measurement requires that the thing being measured stays still. You need measured quality before you expand, because expansion without a baseline means you cannot tell whether the second use case degraded the first.
The mistake that costs the most is expanding on enthusiasm rather than evidence. The tender agent works well, so someone asks whether it could also handle supplier contracts. It probably could, eventually. But if the eval set only covers tenders, the first contract regression will be discovered by a user, and the credibility cost of that is much higher than the delay would have been.
Expand when you can answer three questions: what does quality look like in the new area, do we have examples of it, and will the existing eval set still pass afterwards.
What good looks like
One workflow, one domain, and a definition of done that a domain expert agrees with.
Enough real examples that the common traps are represented, not just the clean cases.
A quality measure that existed before the expansion, so that regression is visible.
An honest scope statement. "This handles South African ICT tenders" is a stronger claim than "this handles tenders", because the first one can be true.
We have written separately about how we built TenderCity and about starting small on tender automation.
Practical checklist
- Name the single workflow, in the language the people doing it use
- Write down the domain rules a newcomer would get wrong in their first month
- Collect real examples, including the messy ones
- Define what a correct output looks like before building
- Measure quality on that one workflow until the number is stable
- Only add the second use case when the first one's evals still pass
How CloudNala can help
We tend to argue for a narrower first scope than clients expect, and the argument is always the same: a narrow agent that a team trusts is worth more than a broad one they check by hand. In practice the first engagement is usually one workflow, the domain rules written down properly, a set of real examples assembled with the people who do the work, and a quality measure that exists before anything is expanded.
Work with CloudNala
CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.
Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.
Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za