← Back to all insights

AI Strategy Published · 10 July 2026

Why AI pilots don't scale

The MIT study that found 95% of pilots move nothing on the P&L points at something specific: the failure is almost never the model. It is the stretch between the demo and the process.

6 min read

A pilot that works and never reaches production is more expensive than a pilot that fails fast. It burns budget, burns internal credibility, and leaves the organisation convinced that “AI isn’t for us”.

And it is the most common outcome.

The number

The GenAI Divide, from MIT’s NANDA project, analysed 300 public deployments, interviewed 52 executives and surveyed 153 leaders. Its conclusion: despite $30–40 billion of enterprise investment, 95% of generative AI projects produced no measurable return.

And the cause it names is not technical: data readiness, workflow integration and the absence of a defined outcome.

The five jumps a pilot doesn’t make

A pilot proves something is possible. Production demands five other things, and each one is where projects die:

1. From the nice case to the real one. The pilot is tested on twenty clean documents somebody picked. Production brings the skewed scan, the supplier who changed format and the customer who writes in three languages. The demo’s 94% accuracy becomes 71%.

2. From one person to a team. In the pilot it is used by whoever built it, who knows what to ask. In production it is used by twelve people who weren’t in the meetings and whom nobody told about the limits.

3. From enthusiasm to maintenance. The pilot has a motivated owner. Production needs someone on call when it breaks at seven on a Tuesday evening, and an annual budget nobody forecast.

4. From output to process. This is the big one. The pilot produces an output; the process needs to know what happens to it. Who reviews it, what happens when it’s wrong, where it is stored, which system consumes it next.

5. From “it works” to “we know it works”. Without continuous measurement, nobody can claim the system still performs in six months. And without that claim, nobody signs the expansion.

What the ones who scale do differently

The study itself identifies a 5% that does extract value. It is not distinguished by technology, but by integrating AI into a specific workflow rather than deploying a generic tool.

What we see in the ones who manage it:

  • The pilot is designed like production from day one, just at lower volume. Same dirty data, same real users, same odd cases.
  • There is a process owner, not a project owner. Someone whose job gets better or worse depending on the result.
  • The success criterion is written before starting, with a number and a date.
  • Operations are budgeted, not just the build. Like any other system.
  • The scope is uncomfortably small. One process, one team, one metric.

The kickoff question

Before approving a pilot, this conversation saves months:

“If the pilot goes perfectly, what has to happen to put it in production, who does it, and on what budget?”

If there is no answer, you are not approving a pilot. You are approving an expensive demo.

Sources

Next step

How ready is your business for AI?

Evaluate your AI maturity in 5 minutes and get free personalised recommendations.

Ready to move beyond the hype?