Most enterprise AI initiatives don't fail at the technology stage. They fail months earlier, in a meeting room, when someone picks the wrong process to start with. The model works, the vendor delivers, the dashboard looks great. Six months later there is still no number a CFO would accept as proof.

This guide is a practical method for choosing a first AI pilot process: four criteria and a scoring exercise you can run in a single leadership meeting. It also covers the traps we keep seeing capable teams walk into.

The real failure point is selection, not technology

MIT's much-quoted 2025 research put a number on it: roughly 95% of enterprise generative AI pilots produce no measurable P&L impact. The instinctive explanation, immature technology, doesn't hold up. Document extraction, conversational AI, and process automation have been running in production for years across thousands of deployments.

What separates successful pilots from failed ones is the shape of the process they were pointed at. A first pilot has one job: produce an undeniable, measured result inside one quarter. Every selection decision should serve that goal, not strategic ambition and not whatever a board member read about over the weekend.

"Your first pilot is not a technology test. It is a credibility test. Pick the process that makes the result impossible to argue with."

The four selection criteria

1. Data availability

The process must already leave a digital trail — event logs in your ERP, tickets in your service desk, documents in a defined inbox. If the data has to be created before the pilot can begin, you are running a data project with an AI pilot attached to it. That is a valid project, but it will not deliver results in 90 days.

2. Frequency and volume

A process that runs 500 times a day produces statistically meaningful results in two weeks. A process that runs twice a month needs a year to prove anything. High frequency gives you fast feedback and a result you can defend with sample sizes instead of anecdotes.

3. Visible pain

Choose a process whose cost is already on a slide somewhere: backlog hours, processing time per case, error rates, SLA breaches. If leadership already tracks the metric, you don't have to convince anyone the problem matters. You only have to move the number.

4. Contained scope

One process owner. One system boundary. One handover or fewer. Every additional department in the loop multiplies coordination cost and gives the pilot more ways to stall for reasons that have nothing to do with AI.

A scoring framework you can run in one meeting

Score each candidate process from 1 (low) to 3 (high) on each criterion. Anything scoring 10 or above is a strong pilot candidate. Anything below 7 should wait, no matter how strategically exciting it sounds.

Criterion Invoice Processing Customer Onboarding Redesign
Data availability3 — invoices already digital, ERP event logs exist1 — journey spans email, calls, three systems
Frequency & volume3 — hundreds per day2 — dozens per week
Visible pain3 — processing cost per invoice already reported1 — "experience" pain, no agreed metric
Contained scope2 — AP team owns it, one approval handover1 — sales, legal, finance, and IT all involved
Total11 / 12 — strong pilot5 / 12 — not yet

Notice what the framework does: it removes opinion from the room. The onboarding redesign may well be the more important initiative long term. As a first pilot, though, it would burn your credibility before producing a single defensible number.

Three traps that sink first pilots

What a well-selected pilot looks like

One of our e-commerce deployments illustrates the pattern. Order processing scored high on every criterion: fully digital order data, hundreds of cases daily, processing time already tracked by operations, and a single owning team. The pilot went live in six weeks. Within the quarter, 80% of manual order touchpoints were gone and batch processing time fell from four hours to under twelve minutes. Nobody in the building argued with those numbers.

The pattern to copy: the pilot's success was decided before any technology was deployed. The process was selected so that the data existed, the volume guaranteed fast proof, the metric was already on a management dashboard, and one team could say yes to everything.

A first win like that buys you the budget and the political capital to take on the harder, more strategic processes later. Including, eventually, that onboarding redesign.

Frequently asked questions

The one that scores highest on four criteria: existing digital data, high daily frequency, a pain metric leadership already tracks, and contained scope with a single owner. In most enterprises that points to high-volume document-driven processes — invoice processing, order processing, claims intake.

Live, measured results within 90 days. If a proposed pilot can't credibly deliver inside a quarter, the scope is too broad or the data isn't ready — fix that before starting, not during.

They fail at selection, before deployment: politically complex processes, no digital data trail, or no success metric agreed upfront. The technology is rarely the binding constraint.

Getting started

You don't have to build the candidate list from scratch. We maintain a ranked library of 30 high-impact processes across 7 business functions (sales, finance, operations, service, HR, IT, and marketing), each scored by pilot ROI, with the data sources, process owners, and break points already mapped. Pick one from the list and we'll scope a fixed-scope pilot in a single call, with live results in under 90 days.