AI & Data27 August 20266 min read

No AI Without Information Architecture — Part 7 of 10

Most of the data you collect is never used

You pay for the storage, the pipelines, and the field time of everyone filling in the forms. Collection without a decision attached to it is not a data strategy — it is an expensive filing habit.

#AI & Data#Digital Transformation#Public Sector#Solution Architecture

The figure people quote is that around 70% of collected data is never used for anything. I have no way to verify that number and I would treat any precise version of it with suspicion.

What I would not dispute is the direction. Every organisation I have looked at closely is collecting substantially more than it uses, and paying for all of it.

Collected, stored, paid for, never openedWHAT IS COLLECTEDEvery form, every field, every monthWHAT INFORMS A DECISIONA fraction of itCollected anyway, stored anyway, paid for anywayCollection without a decision attached to it is not a data strategy,it is an expensive filing habit
The commonly cited figure is that most collected data is never used for anything. Whatever the exact proportion, the cost is not hypothetical: it is paid every month, in storage, in pipelines, and in the field time of the people filling the forms in.

The cost is not hypothetical

It is easy to shrug at unused data because storage is cheap. Storage is the smallest part of the bill.

Field time. Someone fills in every field on every form. If a form has forty fields and eleven of them inform a decision, you are spending the other twenty-nine across every submission, by every worker, every cycle. In a programme with staff across a province, that is an enormous recurring cost paid in the time of people whose actual job is something else.

Pipeline maintenance. Every collected field has to be ingested, validated, transformed and kept working when the source changes. Unused fields break as often as used ones and nobody notices for months, which is its own kind of problem.

Quality drag. Long forms are filled in worse. Ask for twenty-nine unnecessary things and the eleven that matter get less care. Unused collection actively degrades the data you do use.

Sensitivity you did not need. Every unnecessary personal field is an obligation you have taken on — retention, access control, breach exposure — for information nothing consumes. Under POPIA, collecting more than is necessary for the stated purpose is not merely wasteful; it is a compliance problem you created for no benefit.

Why it happens

Nobody decides to collect useless data. Two entirely reasonable instincts produce it.

The first is collect it while we can. Getting a form changed is administratively painful, so when a form is open, everyone adds their field just in case. Every addition is individually justified. The aggregate has never been reviewed.

The second is data is an asset. This has been said for a decade and it is half-true in a damaging way. Data is an asset in the way that inventory is an asset — valuable if it moves, a cost if it sits. An organisation that has internalised "collect everything, value will emerge" has adopted a strategy with no completion condition and no way to fail visibly.

Work backwards instead

The question that comes before the form is designedWORKING FORWARDS — WHAT MOST PROGRAMMES DOCapture whatis easyStore all of itLook for valueafterwardsRarely find itWORKING BACKWARDSName thedecisionName whomakes itWork out whatthey need to seeCollect thatA data strategy is a list of decisions you intend to improve, not a list of sources you intend to ingest
Working backwards from a decision produces a shorter form, less storage, and data somebody is waiting for. Working forwards from what is easy to capture produces volume and no answers.

The top row is what most programmes do, and it has a fatal property: there is no point at which anyone is accountable for value, because value was always going to come later.

The bottom row starts from the other end. Name the decision you want to improve. Name the person who makes it. Work out what they would need to see to make it better. Collect that.

It produces a shorter form, less storage, less sensitive data, and — this is the part that matters — someone who is actually waiting for the output. Data with a waiting consumer gets used, gets quality-checked, and gets defended at budget time. Data collected speculatively has none of those.

The honest counter-argument

There is a real objection here and it deserves a straight answer rather than a dismissal.

Some of the most valuable analysis comes from data nobody knew they would need. Work backwards too rigidly and you optimise away the raw material for questions that have not been asked yet. That is a genuine tension, not a rhetorical one.

Where I land: be strict about what you require people to enter, and relaxed about what you retain from what systems already emit.

Asking a field worker for a field costs real human time and degrades everything around it — that bar should be high, and "we might need it" does not clear it. Retaining an event log a system produces anyway costs almost nothing and preserves optionality — that bar can be low.

The waste is concentrated almost entirely in the first category, and so is the compliance exposure. Most organisations have it backwards: long forms and short log retention.

A data strategy is a list of decisions

The most useful reframe I can offer.

A data strategy is often written as an inventory of sources to ingest and a platform to ingest them into. That document is not a strategy; it is a shopping list, and it can be executed completely while changing nothing about how the organisation is run.

A real data strategy is a list of decisions you intend to improve, with the person who makes each one named, and what they would need in order to make it better. Everything else — architecture, platform, governance, sequencing — follows from that list and can be judged against it.

It also gives you a way to say no. When someone proposes a new collection point, the question is: which decision on the list does this improve? If none, it goes on a backlog rather than into the form.

That question alone will halve most collection programmes, and the half you keep will be the half anyone uses.

How CloudNala can help

The exercise we run early on most data engagements is deliberately unglamorous: name the decisions worth improving, name who makes them, and work backwards to what needs collecting. It usually results in us building less than the client expected, and in the data that does get collected having someone waiting for it — which is the difference between a programme that survives its first budget review and one that does not.


Work with CloudNala

CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.

Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.

Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za