Independent AI project review — serving clients nationwide Mon–Fri 9am–5pm CST  |  +1 (414) 635-2154
July 14, 2026 · Research

The 95% Problem: Why Most AI Pilots Never Reach the P&L

MIT's NANDA initiative published a number this year that every board in the country should have read twice: 95% of enterprise generative AI pilots deliver no measurable profit-and-loss impact. Vendors reacted the way vendors react to inconvenient data — they argued about methodology. We think that's the wrong argument.

We've sat in enough project reviews to tell you the number is directionally right, and that it has almost nothing to do with model quality. In every stalled project we've examined, the model worked well enough to be useful. What was missing was someone whose job it was to prove that.

Nobody owns the verdict

A typical pilot gets funded by an innovation budget, championed by whoever brought the vendor in, and evaluated by the same people who chose it. There is no independent checkpoint where someone asks, plainly: did this move a number that shows up on a financial statement? Engagement metrics get substituted for P&L metrics because engagement metrics are easier to produce and harder to argue with.

That's not a technology failure. It's an org-design failure, and it's the same failure a CFO would flag in any other capital project. Nobody would let a plant expansion run for eighteen months without a capital-appropriation review. AI pilots run that long routinely, on vibes.

What changes the number

In the handful of pilots we've seen actually convert to P&L impact, three things were true that weren't true in the other 95%:

  • A single named metric was chosen before the pilot started, not after it looked good.
  • Someone outside the buying decision was accountable for saying, in writing, whether the metric was hit.
  • There was a pre-agreed exit point if it wasn't — not an indefinite extension.
The question was never “does the AI work.” It was “did anyone ever put a number on what working would mean, and check it.”

That's the entire premise behind the AI Reality Check: an independent, written verdict — continue, restructure, replace, or terminate — on a single stalled or underperforming project, in five days, from someone with no stake in the original purchase decision.

The 95% statistic isn't a reason to stop trying. It's a reason to stop grading your own homework.

See how the Reality Check works →