← Back to all insights

Data & Analytics Published · 3 July 2026

Data quality is the ceiling, and no model raises it

MIT identified data readiness as one of the three reasons 95% of pilots deliver nothing. It is not a problem a better model fixes, because it is not a model problem.

6 min read

There is a conversation that repeats: the system underperforms, the team proposes switching models, they switch, and it performs about the same. Then someone proposes fine-tuning on in-house data. Still the same.

The problem was almost always upstream.

What the evidence says

The GenAI Divide from MIT, explaining why 95% of pilots produce no return, names three causes and the first is data readiness. The failure, they write, is almost never the model.

It matches what we find in every pre-project audit we run.

The five ways data limits you

1. The data you need doesn’t exist. You want to predict which customer will churn, but nobody ever recorded why any of them left. No model learns from an empty field.

2. It exists, but everyone fills it in differently. Three salespeople, three definitions of “active customer”. The model learns the noise in your internal processes, not the pattern in your business.

3. It is spread across systems that don’t talk. The ERP knows what was invoiced, the CRM what was promised, and an operations spreadsheet what actually happened. None has the whole story.

4. History reflects decisions, not reality. If credit was approved on a biased criterion for five years, a model trained on that history will reproduce the bias with industrial efficiency.

5. Nobody owns fixing it. The data belongs to everyone, which means nobody. When an inconsistency appears, it gets documented as an exception and life goes on.

What has changed (and what hasn’t)

Something important: current models tolerate mess far better than they did three years ago. Free text, inconsistent formats, scanned documents — things that once demanded a cleanup project beforehand are now processed reasonably well.

What has not changed: if the data doesn’t exist, or is systematically wrong, there is still nothing to be done. Tolerance for noise went up; the ability to invent information that was never recorded did not.

That distinction is what lets you decide well. Plenty of projects shelved because “we need to clean the data first” are viable today with no cleaning at all.

The two-week audit

Before quoting any implementation we run a short data audit. Not for academic rigour: because it saves the client money.

What it looks at:

  • Does the data the use case needs exist? And if so, since when and with what coverage?
  • How much variation is there in how it is recorded? Measured on a sample, not by opinion.
  • What percentage of records is complete in the fields that matter?
  • Is there a ground truth to measure against? Without one there is no evaluation possible.
  • Who can fix the source if needed, and will they want to?

Two weeks of this avoids six months of useless modelling. It is the most boring recommendation we give and the one we have been thanked for most.

The practical rule

If you can’t explain how the data is generated, don’t build on top of it. And if the process generating it is broken, fix that first: it is cheaper than any model trying to compensate, and it improves the business on its own.

Sources

Next step

How ready is your business for AI?

Evaluate your AI maturity in 5 minutes and get free personalised recommendations.

Ready to move beyond the hype?