The figure people quote is that around 70% of collected data is never used for anything. I have no way to verify that number and I would treat any precise version of it with suspicion.
What I would not dispute is the direction. Every organisation I have looked at closely is collecting substantially more than it uses, and paying for all of it.
The cost is not hypothetical
It is easy to shrug at unused data because storage is cheap. Storage is the smallest part of the bill.
Field time. Someone fills in every field on every form. If a form has forty fields and eleven of them inform a decision, you are spending the other twenty-nine across every submission, by every worker, every cycle. In a programme with staff across a province, that is an enormous recurring cost paid in the time of people whose actual job is something else.
Pipeline maintenance. Every collected field has to be ingested, validated, transformed and kept working when the source changes. Unused fields break as often as used ones and nobody notices for months, which is its own kind of problem.
Quality drag. Long forms are filled in worse. Ask for twenty-nine unnecessary things and the eleven that matter get less care. Unused collection actively degrades the data you do use.
Sensitivity you did not need. Every unnecessary personal field is an obligation you have taken on — retention, access control, breach exposure — for information nothing consumes. Under POPIA, collecting more than is necessary for the stated purpose is not merely wasteful; it is a compliance problem you created for no benefit.
Why it happens
Nobody decides to collect useless data. Two entirely reasonable instincts produce it.
The first is collect it while we can. Getting a form changed is administratively painful, so when a form is open, everyone adds their field just in case. Every addition is individually justified. The aggregate has never been reviewed.
The second is data is an asset. This has been said for a decade and it is half-true in a damaging way. Data is an asset in the way that inventory is an asset — valuable if it moves, a cost if it sits. An organisation that has internalised "collect everything, value will emerge" has adopted a strategy with no completion condition and no way to fail visibly.
Work backwards instead
The top row is what most programmes do, and it has a fatal property: there is no point at which anyone is accountable for value, because value was always going to come later.
The bottom row starts from the other end. Name the decision you want to improve. Name the person who makes it. Work out what they would need to see to make it better. Collect that.
It produces a shorter form, less storage, less sensitive data, and — this is the part that matters — someone who is actually waiting for the output. Data with a waiting consumer gets used, gets quality-checked, and gets defended at budget time. Data collected speculatively has none of those.
The honest counter-argument
There is a real objection here and it deserves a straight answer rather than a dismissal.
Some of the most valuable analysis comes from data nobody knew they would need. Work backwards too rigidly and you optimise away the raw material for questions that have not been asked yet. That is a genuine tension, not a rhetorical one.
Where I land: be strict about what you require people to enter, and relaxed about what you retain from what systems already emit.
Asking a field worker for a field costs real human time and degrades everything around it — that bar should be high, and "we might need it" does not clear it. Retaining an event log a system produces anyway costs almost nothing and preserves optionality — that bar can be low.
The waste is concentrated almost entirely in the first category, and so is the compliance exposure. Most organisations have it backwards: long forms and short log retention.
A data strategy is a list of decisions
The most useful reframe I can offer.
A data strategy is often written as an inventory of sources to ingest and a platform to ingest them into. That document is not a strategy; it is a shopping list, and it can be executed completely while changing nothing about how the organisation is run.
A real data strategy is a list of decisions you intend to improve, with the person who makes each one named, and what they would need in order to make it better. Everything else — architecture, platform, governance, sequencing — follows from that list and can be judged against it.
It also gives you a way to say no. When someone proposes a new collection point, the question is: which decision on the list does this improve? If none, it goes on a backlog rather than into the form.
That question alone will halve most collection programmes, and the half you keep will be the half anyone uses.
How CloudNala can help
The exercise we run early on most data engagements is deliberately unglamorous: name the decisions worth improving, name who makes them, and work backwards to what needs collecting. It usually results in us building less than the client expected, and in the data that does get collected having someone waiting for it — which is the difference between a programme that survives its first budget review and one that does not.
Work with CloudNala
CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.
Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.
Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za