AI Strategy27 August 20267 min read

No AI Without Information Architecture — Part 6 of 10

The semantic layer is your hallucination control

An agent asked for revenue returns a number either way. Whether it matches what the finance director means by revenue depends on whether a definition exists somewhere the agent can read.

#AI Strategy#AI Agents#AI Governance#AI & Data

The business glossary has been the least fashionable artefact in data management for twenty years. Every methodology recommends one. Almost nobody maintains one, because the payoff was always diffuse — fewer arguments in meetings, roughly.

Agents have changed the economics of that completely, and it is the most under-discussed consequence of pointing AI at organisational data.

The layer that decides whether the answer is trueA questionin plain EnglishThe agentand its toolsSemantic layerdefinitions,glossary, lineageThe dataWITH THE LAYEROne agreed definition of revenueWITHOUT ITA confident number nobody can defend
An agent asked for revenue will return a number either way. Whether that number matches what the finance director means by revenue depends entirely on whether a definition exists somewhere it can read.

The revenue problem

Ask an agent with database access: what was revenue last quarter?

It will answer. It will find something plausibly named, sum it, and present a figure in a confident sentence. The figure will be wrong roughly as often as the organisation has more than one reasonable definition of revenue — which is to say, most of the time.

Does it include VAT? Does it net off credits? Does it recognise on invoice or on payment? Does it include intercompany? Which of the three systems is authoritative when they disagree?

A human analyst knows to ask. They have been corrected before. An agent has table names, and table names do not carry that.

The same problem, in plainer clothes: what was our on-time delivery rate? Does on-time include weekends? Is the clock from order or from dispatch? Is a partial delivery on time? Every one of those is a real definitional choice, every organisation has made them somewhere, and almost none have written them down anywhere a machine can read.

This is not the model being unreliable. The model is doing precisely what was asked with the information available. The information was incomplete.

What actually sits underneath a plain-English question

The pitch for natural-language analytics is that anyone can now ask anything. What that skips is everything the question needs underneath it.

What a plain-English question still needs underneath it"How many did we complete in the Northern region last quarter?"Which table holds this, and how does it join to the others?What does this business term actually mean here?Is this person allowed to see these rows and columns?Where did this number come from, and can I show my working?How much is this query about to cost?
A connection to the data answers none of these. They are answered by the catalogue, the glossary, the access model, the lineage and the cost controls — which is the governance work, arriving under a different name.

Look at that list. Not one of those is answered by connecting a model to a database.

They are answered by the catalogue (what exists and how it relates), the glossary (what these terms mean here), the access model (who may see which rows and columns), lineage (where this number came from and what was done to it), and cost controls (what this query is about to spend).

That is the governance work. It has simply arrived under a new name, with a much more concrete payoff than it has ever had before. For twenty years the argument for a business glossary was that it reduces confusion. The argument now is that without it, your AI produces confident wrong answers at machine speed — and people believe them, because the prose is fluent and there is no hesitation in it.

That is a materially easier case to make to an executive.

Knowledge graphs, and why the catalogue earns its keep

There is a second-order benefit that is worth understanding, because it changes how much the catalogue is worth.

A well-maintained catalogue does not only record what exists. It records relationships — this table joins to that one on this key, this measure derives from those sources, this field is a restatement of that one. That relationship map is a knowledge graph, whether or not anyone calls it that.

For an agent, that graph is the difference between guessing at joins and navigating a described structure. Give a model five table names and it will infer relationships from column naming, which works until two tables both have a column called region_id that mean different things. Give it the graph and it does not have to guess.

So the catalogue stops being documentation and becomes runtime infrastructure. That is a genuine change in what it is for, and it is why "we will document it later" is now a much more expensive decision than it used to be.

The uncomfortable implication

If the semantic layer determines whether AI answers are trustworthy, then AI readiness is mostly a definitional problem, not a technical one.

And definitional problems are organisational. Deciding what revenue means is not a data engineering task — it requires finance, operations and whoever owns the number to agree, and sometimes they disagree for good reasons that have never had to be resolved because everyone quietly used their own version.

The agent forces the resolution. It cannot hold two definitions and use the right one contextually. Something has to be written down.

This is genuinely hard and it is where these programmes slow down. It is also, I would argue, most of the value — an organisation that has agreed what its terms mean is better run regardless of what technology it deploys. The AI is the forcing function, not the benefit.

What a connection does not give you

One correction worth making explicitly, because the tooling conversation moves fast and the distinction gets lost.

The protocols and connectors that let a model reach a database are connection mechanisms. They are useful and they solve a real integration problem. They do not:

  • make the returned answer correct
  • enforce who is allowed to see what
  • record who asked what, when
  • prevent a vague question from scanning everything you own
  • ground a term in an agreed definition

Every one of those is a separate piece of work, and each is the kind of thing that is invisible in a demo and unavoidable in production. A demo has one user, who is you, asking questions you chose, against data you know. Production has hundreds of users asking things you did not anticipate, and the gap between those is filled with exactly the five items above — plus the cost question, which is its own article.

Where to start

If you are earlier in this than you would like, the useful first move is small: take the ten terms that appear most often in your executive reporting and write down what each one means, precisely, with the rule for edge cases and the authoritative source.

Ten definitions. It is not a governance programme, it takes an afternoon of the right people's time, and it will surface at least one disagreement that has been costing you quietly for years.

Everything else in the semantic layer builds outward from that.

How CloudNala can help

The work we do here is less technical than clients expect. Most of it is getting the right people in a room to agree what a term means and then writing it somewhere durable and machine-readable — so that both the reporting and, later, any agent, are grounded in the same definitions. It is unglamorous and it is the single strongest predictor we have seen of whether an AI deployment holds up once real users arrive.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za