Why AI pilots fail: 95% deliver nothing — and where the returns actually are
Why AI pilots fail is now well documented. Where the returns actually sit, and how to pick a first project that proves itself inside one quarter.
Most AI pilots fail because the work was chosen badly and never measured — not because the models are weak. MIT's Project NANDA reviewed enterprise adoption and found that about 95% of generative AI pilots produced no measurable return against 30 to 40 billion dollars of investment. The same research located the highest returns in back-office processes, while 50 to 70% of budgets went into sales and marketing, where returns were lowest.
Why do most AI pilots fail?
They fail on fit and measurement, not intelligence. The MIT report's diagnosis is that the tools "don't adapt, don't retain feedback and don't fit daily workflows" — the capability was adequate and the integration was not.
The numbers worth carrying into a meeting:
- 95% of pilots produced no measurable return, across more than 300 disclosed initiatives, 52 interviews and 153 executive surveys — MIT Project NANDA, 2025
- 50–70% of AI budgets went to sales and marketing pilots, where ROI was lowest — same study
- Back-office automation produced the highest returns, by cutting processing, reconciliation and outsourcing costs — same study
Read the definition before reacting to the headline. Success meant getting past the pilot stage with measurable results six months on. That is a demanding bar, and some analysts argue it misses value that never becomes a KPI. It is also the bar your finance director will use.
Where do the returns actually come from?
From repetitive back-office work, for a structural reason. Sales and marketing pilots are easy to start and almost impossible to judge. A generated email sequence looks impressive in a meeting; six months later nobody can separate its effect from seasonality, pricing or a new hire. The pilot cannot fail clearly, so it cannot succeed clearly either.
Back-office work is the opposite. It is repetitive, follows an existing procedure, uses known sources, and usually has a right answer. When you automate part of it you can measure before and after, because the before was already being measured — in hours, in errors, in how late the report went out.
Is high output the same as a result?
No, and conflating them is how pilots produce activity for months without value. Researchers at Carnegie Mellon staffed a simulated software company entirely with AI agents and found they struggled with basic office tasks. In one widely discussed demonstration, a browser agent applied to 100 jobs in 40 minutes and produced zero interviews.
The question is never how much the system produced. It is how much finished work left the building without a person redoing it — a distinction covered in the time saving that verification eats.
What does a measurable AI pilot need?
Six things fixed before anything is built:
One repeated process. Not a category — one process. "Supplier invoice reconciliation" is a pilot. "Finance" is not.
One owner. A named person who already owns the outcome and can say whether the result is acceptable.
Defined sources. Which files, which mailbox, which system, which fields. If the answer is "it depends who does it that week", you have found something to fix before any AI is involved.
A before baseline. How long it takes now, how often it is late, how often it is wrong. Measure it for a few cycles first. Almost nobody does, which is why almost nobody can prove a result afterwards.
A definition of finished. The artefact in the real template, sources named, exceptions flagged. If a person still has to assemble it, the saving will move rather than appear.
An approval point. Decide in advance which step a person must sign — usually the first irreversible one.
This is where tooling either helps or gets in the way. Kvantia Harness is a supervision layer for an AI worker on a Windows PC your business controls: capabilities are off by default, irreversible actions wait for a person, credentials stay out of the model context, and every action and source is recorded. That last property is what makes a pilot assessable at all — you can answer "did it work" from the run record instead of from opinion. How it works walks through a routine end to end.
Which questions end a pilot early?
Three, and early is the cheapest time for a pilot to end.
Is the process written down, or does it live in one person's head? If it is in someone's head, the first output of the project is documentation — valuable, but not yet an AI project.
Are the source numbers trustworthy? Automation makes bad inputs travel faster and look more official.
Would you let a new employee do this in week one, with review? If not, the reason is usually that consequences are severe and checks are informal. That is a signal to automate the preparation and keep the decision.
Where does that leave your shortlist?
On the recurring, procedural, verifiable work that already has an owner and a deadline: the weekly management pack, the register that must match a source system, the reconciliation nobody enjoys, the follow-ups that are always late by Thursday.
That work is dull, which is exactly why it qualifies. Dull means repeated, repeated means measurable, and measurable means you will know within two months whether it was worth doing — which is the thing 95% of pilots could not say. For the selection test itself, see which tasks to automate with AI first.