Most enterprise AI initiatives don't fail at the technology stage. They fail months earlier, in a meeting room, when someone picks the wrong process to start with. The model works, the vendor delivers, the dashboard looks great. Six months later there is still no number a CFO would accept as proof.
This guide is a practical method for choosing a first AI pilot process: four criteria and a scoring exercise you can run in a single leadership meeting. It also covers the traps we keep seeing capable teams walk into.
The real failure point is selection, not technology
MIT's much-quoted 2025 research put a number on it: roughly 95% of enterprise generative AI pilots produce no measurable P&L impact. The instinctive explanation, immature technology, doesn't hold up. Document extraction, conversational AI, and process automation have been running in production for years across thousands of deployments.
What separates successful pilots from failed ones is the shape of the process they were pointed at. A first pilot has one job: produce an undeniable, measured result inside one quarter. Every selection decision should serve that goal, not strategic ambition and not whatever a board member read about over the weekend.
"Your first pilot is not a technology test. It is a credibility test. Pick the process that makes the result impossible to argue with."
The four selection criteria
1. Data availability
The process must already leave a digital trail — event logs in your ERP, tickets in your service desk, documents in a defined inbox. If the data has to be created before the pilot can begin, you are running a data project with an AI pilot attached to it. That is a valid project, but it will not deliver results in 90 days.
2. Frequency and volume
A process that runs 500 times a day produces statistically meaningful results in two weeks. A process that runs twice a month needs a year to prove anything. High frequency gives you fast feedback and a result you can defend with sample sizes instead of anecdotes.
3. Visible pain
Choose a process whose cost is already on a slide somewhere: backlog hours, processing time per case, error rates, SLA breaches. If leadership already tracks the metric, you don't have to convince anyone the problem matters. You only have to move the number.
4. Contained scope
One process owner. One system boundary. One handover or fewer. Every additional department in the loop multiplies coordination cost and gives the pilot more ways to stall for reasons that have nothing to do with AI.
A scoring framework you can run in one meeting
Score each candidate process from 1 (low) to 3 (high) on each criterion. Anything scoring 10 or above is a strong pilot candidate. Anything below 7 should wait, no matter how strategically exciting it sounds.
| Criterion | Invoice Processing | Customer Onboarding Redesign |
|---|---|---|
| Data availability | 3 — invoices already digital, ERP event logs exist | 1 — journey spans email, calls, three systems |
| Frequency & volume | 3 — hundreds per day | 2 — dozens per week |
| Visible pain | 3 — processing cost per invoice already reported | 1 — "experience" pain, no agreed metric |
| Contained scope | 2 — AP team owns it, one approval handover | 1 — sales, legal, finance, and IT all involved |
| Total | 11 / 12 — strong pilot | 5 / 12 — not yet |
Notice what the framework does: it removes opinion from the room. The onboarding redesign may well be the more important initiative long term. As a first pilot, though, it would burn your credibility before producing a single defensible number.
Three traps that sink first pilots
- The most painful process. Maximum pain usually means maximum politics. The most broken process in the company is broken for organisational reasons, and an AI pilot will inherit every one of them.
- The executive pet idea. A use case chosen for who proposed it, not how it scores. If it can't pass the four criteria, park it respectfully and let the scoring matrix take the blame.
- The process with no digital footprint. If the work happens in hallway conversations and spreadsheets on personal drives, there is nothing for AI to learn from yet. Digitise first, pilot second.
What a well-selected pilot looks like
One of our e-commerce deployments illustrates the pattern. Order processing scored high on every criterion: fully digital order data, hundreds of cases daily, processing time already tracked by operations, and a single owning team. The pilot went live in six weeks. Within the quarter, 80% of manual order touchpoints were gone and batch processing time fell from four hours to under twelve minutes. Nobody in the building argued with those numbers.
The pattern to copy: the pilot's success was decided before any technology was deployed. The process was selected so that the data existed, the volume guaranteed fast proof, the metric was already on a management dashboard, and one team could say yes to everything.
A first win like that buys you the budget and the political capital to take on the harder, more strategic processes later. Including, eventually, that onboarding redesign.
Frequently asked questions
The one that scores highest on four criteria: existing digital data, high daily frequency, a pain metric leadership already tracks, and contained scope with a single owner. In most enterprises that points to high-volume document-driven processes — invoice processing, order processing, claims intake.
Live, measured results within 90 days. If a proposed pilot can't credibly deliver inside a quarter, the scope is too broad or the data isn't ready — fix that before starting, not during.
They fail at selection, before deployment: politically complex processes, no digital data trail, or no success metric agreed upfront. The technology is rarely the binding constraint.
Getting started
You don't have to build the candidate list from scratch. We maintain a ranked library of 30 high-impact processes across 7 business functions (sales, finance, operations, service, HR, IT, and marketing), each scored by pilot ROI, with the data sources, process owners, and break points already mapped. Pick one from the list and we'll scope a fixed-scope pilot in a single call, with live results in under 90 days.