You've probably seen the number. MIT's 2025 research on enterprise AI found that roughly 95% of generative AI pilots produce no measurable P&L impact. It gets quoted in board meetings, usually right before someone's AI budget gets cut.
Here's the thing: the skepticism is justified, but the conclusion most people draw from it is wrong. The pilots aren't failing because the technology doesn't work. They're failing because of how they're set up. After watching this play out across dozens of deployments, the failure patterns are so consistent you can spot them in the kickoff meeting.
The five failure patterns
1. Open-ended scope
The pilot brief says something like "explore what AI can do for operations." That isn't a scope, it's a research grant. Without a boundary, the project expands until it touches every system and stakeholder in the company, and then it stalls.
2. No success metric agreed before kickoff
If the team can't say, in one sentence, what number will prove the pilot worked, the pilot has already failed. Defining success after the results come in guarantees an argument instead of a decision.
3. Built on data that doesn't exist yet
The use case sounds great in the deck. Then week three arrives and someone discovers the "historical data" is a folder of inconsistent spreadsheets on a shared drive. Now the AI pilot is secretly a six-month data cleanup project, and nobody budgeted for that.
4. No single accountable owner
When a pilot belongs to a committee, every decision takes two weeks. The successful pilots we've seen have one named person who can approve changes, unblock access, and answer questions the same day.
5. No exit criteria
Some vendors are happy to let a pilot run forever. Monthly check-ins, encouraging demos, no decision point. If the engagement doesn't have a date where everyone looks at the numbers and decides go or no-go, it isn't a pilot. It's a subscription.
"If the engagement doesn't have a date where everyone looks at the numbers and decides go or no-go, it isn't a pilot. It's a subscription."
What the successful 5% do differently
The pilots that make it into production share a recognisable shape. They pick one process, not one department. The success metric is agreed in writing before kickoff, and it's a number leadership already tracks. The scope is fixed, so the project can't quietly grow. The whole thing is time-boxed to a quarter. And the path to production is defined on day one, so a successful pilot doesn't end with "great, now what?"
None of this is exotic. It's the same discipline you'd apply to any capital project. The difference with AI is that the hype gives everyone permission to skip the discipline, and the 95% number is what that permission costs.
The anatomy of a 90-day pilot
Here is the structure we run, week by week. The exact split varies with the process, but the phases don't.
| Phase | Weeks | What happens | Output |
|---|---|---|---|
| Scoping | 1–2 | Process boundary fixed, data verified to exist, success metric agreed in writing | Signed one-page pilot charter |
| Build | 3–6 | Connectors deployed, models configured against real historical data | Working system on live data |
| Shadow run | 7–10 | System runs alongside the existing process; outputs compared case by case | Accuracy and throughput evidence |
| Decision | 11–12 | Results measured against the charter metric, in front of the sponsor | Go / no-go with production plan |
The shadow run deserves a comment, because it's the phase most pilots skip. Running the AI system in parallel with the human process, on the same live cases, is what turns a demo into evidence. When the comparison covers four weeks and several thousand cases, the go/no-go meeting is short.
Worth noting: a no-go is a successful outcome too. A pilot that proves a process isn't ready, in twelve weeks and at fixed cost, has saved you from finding that out in production over two years. The failure mode isn't stopping. It's drifting.
Seven questions to ask any AI vendor
Before signing a pilot agreement, put these to the vendor. The answers tell you which side of the 95/5 split you're heading for.
- What is the success metric, in writing? If they resist committing to a number, they don't expect to hit one.
- Where exactly does the scope end? Listen for a process boundary, not a department name.
- Can you show that the required data exists today? Ask them to verify it before kickoff, not during.
- Who owns this on our side? A good vendor insists on a single named owner. A bad one is happy with "the team".
- What happens at the end of the time box? There should be a date, a meeting, and a decision.
- What does the path to production look like? If the pilot succeeds, the next step should already be mapped.
- Can we talk to a customer who ran a pilot of similar scope? Not their biggest customer. A similar one.
Frequently asked questions
The most common causes are open-ended scope, no success metric agreed before kickoff, pilots built on data that doesn't exist yet, no single accountable owner, and no defined exit criteria. The technology itself is rarely the binding constraint.
One quarter. Scope in weeks 1–2, build in weeks 3–6, shadow run in weeks 7–10, and a measured go/no-go decision in weeks 11–12. If a proposed pilot can't deliver inside that window, the scope or the data isn't ready.
Ask for the success metric in writing, the exact scope boundary, proof the data exists today, a single named owner, the decision date at the end of the time box, the path to production, and a reference customer with a pilot of similar scope.
Getting started
If you're planning a first or a second-attempt pilot, start by pressure-testing the process selection itself. Our framework for that is in How to Pick Your First AI Pilot Process. When you've got a candidate, we'll scope a fixed-scope pilot in one call and deliver live, measured results in under 90 days, with the go/no-go date in the calendar from day one.