Skip to content
evvolv.ai
All stories

Essay · 18 June 2026 · 7 min read

The 95% problem is not a model problem

MIT looked at more than 300 enterprise AI deployments and found almost none of them moved the P&L. The interesting part is why, because it is not the part everyone assumes.

In July 2025, MIT's Project NANDA published a study of enterprise generative AI that produced one number everybody repeated and one finding almost nobody did. The number was 95%. The finding was that the failures had very little to do with the models.

The study, The GenAI Divide: State of AI in Business 2025, looked at more than 300 public AI deployments, ran 52 structured interviews with executives, and collected 153 survey responses. Its headline result was that roughly 95% of enterprise generative AI pilots produced no measurable return on the profit and loss statement. Not a small return. No measurable return at all.

Enterprise generative AI pilots by measurable P&L return
5%

Produced a measurable return

95%

Produced none

Source: MIT Project NANDA, The GenAI Divide: State of AI in Business, July 2025

The reflex when you see a number like that is to blame the technology. It is the comfortable conclusion, because it means the problem is somebody else's and it will be fixed by the next release. But the study is fairly direct that this is not what it observed. The models were, for the most part, capable of the tasks being asked of them. What failed was everything around the model.

The gap is organisational, and it is boring

MIT's term for it is the learning gap. A pilot is set up, it demonstrates that the model can do a thing, and then it stops, because demonstrating a capability and installing it into how a company actually operates are separate projects and only the first one is fun.

Some of this is measurement. A pilot that was never given a definition of success cannot produce evidence of it, and a surprising number are commissioned with no agreed metric beyond a general sense that the team should be doing something about AI. Some of it is integration: a system that cannot see the data or reach the tools the job actually requires will produce a convincing demonstration and no work.

And some of it is ownership. When a pilot succeeds, somebody has to change what their team does every day. That person is rarely in the room when the pilot is commissioned, and they have every reason to be unenthusiastic about inheriting it.

Why shadow AI is the most useful data in the report

That last statistic deserves more attention than it got. If official pilots are failing at 95% while unofficial usage is near universal, then the constraint is not appetite, and it is not capability. People want the help and the tools can provide it. What is missing is a version of the help that fits inside the way the organisation actually works, with its approvals, its accountability and its systems of record.

An employee pasting work into a chat window has solved the integration problem by ignoring it. They are the integration. They fetch the context, they judge the output, they carry it back into the real system. It works, at the scale of one person and one task, precisely because a human is doing all the parts that are hard to automate.

The pilot did not fail because the model could not do the task. It failed because nobody could say what should happen next, or who owned it when it did.

What this changes about how you start

If the failure mode is organisational, then the first decision is not which model, it is which job. Specifically: a job with a named owner, a definition of done that two people would agree on, and a real system it has to write back into. Those three constraints eliminate most of the projects that would otherwise become part of the 95%.

  • Pick the most repetitive process, not the most visible one. Visible processes make good demonstrations and bad first projects, because their exceptions are where all the actual work lives.
  • Write down what a good outcome looks like before anything is built. If you cannot, that is the finding, and it is worth more than the pilot would have been.
  • Give it access to the real systems from day one. A pilot that runs on exported spreadsheets is measuring a different job to the one you have.
  • Decide in advance what it is allowed to do without asking. This is the question every operator asks first and most pilots answer last.

None of this is exciting, which is rather the point. The 95% is not evidence that the technology is oversold. It is evidence that the industry has been optimising the demonstration and neglecting the installation, and those are different disciplines.

We build AI coworkers, so we have an obvious interest in the conclusion here. But the reason we build them the way we do, with a named owner, an explicit boundary on what happens without approval, and a write-back into the systems a business already runs, is that we kept watching capable systems fail for reasons that had nothing to do with capability.

Keep reading