Contact
Insights

Systems & AI

Why Your AI Pilot Never Became Part of the Business

A pilot proves a model can do the task. It does not prove your business can absorb the result. Almost everything that decides the second question is left out of the first.

For organisations that ran an AI pilot last year and cannot say what changed.

August 20269 min read

The demonstration went well. That is usually the first sign of trouble.

Nobody declared it dead

The pattern is consistent enough to be worth stating plainly. A pilot is commissioned. Six weeks later there is a demonstration, and it is genuinely impressive — the thing summarises the documents, drafts the replies, reads the invoices. There is a follow-up meeting to discuss rollout. There is a second, quieter one about integration. Then the item stops appearing on the agenda.

Nobody ever decides to stop. The subscription renews once, possibly twice. Eventually someone asks who is still using it, and the answer is two people, occasionally, for something other than the original use case.

The detail that matters is not that it failed. It is that nobody can say what changed. Not what improved and not what got worse — what changed. A programme with a budget, a vendor, a project channel and eleven months of calendar time produced no observable difference in how any work gets done, and the organisation has no account of why.

Why “it wasn’t accurate enough” is usually the wrong answer

Accuracy is the explanation everyone reaches for, because it is measurable, it is nobody’s fault, and it points at the vendor. It is occasionally true. It is much more often a substitute for a harder answer.

Two things are worth separating. The first is that pilot accuracy is not production accuracy, and the reason is rarely the model. Pilots run on inputs that someone selected, cleaned and formatted, usually the person most invested in the pilot succeeding. Production runs on whatever arrives — the scan that is slightly rotated, the email with the attachment missing, the supplier who changed their invoice layout in March. The drop between the two is an input problem wearing an accuracy costume.

The second is that accuracy alone is the wrong test. The question is never “is it accurate?” It is “is it accurate enough to be worth the cost of checking it?” A system that is right eighty-five per cent of the time and gives no signal about which fifteen per cent is wrong requires a human to read one hundred per cent of the output. That system can be very accurate and still add work. Plenty of pilots clear the first bar and quietly fail the second.

A model that is right most of the time, and silent about when it is not, still costs you a full review of everything it produces.

What a pilot is structured to skip

A pilot is a demonstration of capability. An operational system is a change to how work happens. These are different objects, and the distance between them is made of six specific things — none of which demonstrate well, all of which are the actual work:

Read that list again as a list of job descriptions. Every item requires a person in the operating business to change something they own. A pilot, by design, requires nobody to change anything — that is what makes it a pilot, and it is also why so few of them become anything else.

Three shapes of quiet failure

Programmes that stall tend to stall in one of three recognisable ways. Each one looks different from the inside.

The orphan. The system works and has no owner. Nothing announces its decline — a data source changes format, quality drifts, and because nobody is accountable for the output, nobody is watching the output. It is usually discovered months later by a person who assumed it had been checked.

The parallel process. The system produces a good result, and then someone copies that result into the software the business actually runs on. Work has been added, not removed. This one is dangerous precisely because it looks like adoption: usage is high, feedback is positive, and the net effect on hours worked is negative.

The unowned exception. The system handles the routine eighty-five per cent. The remaining fifteen per cent has no defined home, so it returns to the person who previously handled all of it — who now handles the hardest cases only, in smaller volume, having lost the daily context that made those cases tractable. Their work has become harder, and the throughput gain is smaller than anyone modelled.

Instrument

Pilot-to-operation readiness check

Nine questions. Each one must be answered with a name, a system or a number — never with “yes”, “we’ll sort that out” or “the team”. An answer that cannot be made specific is not an answer; it is the point at which this pilot will stop.

  1. 01Whose job changes?Name the role. If no role changes, nothing has been operationalised.
  2. 02What work stops being done?Name the task a person will no longer perform. “They’ll be freed up for higher-value work” is not a task.
  3. 03Where does the input come from in production?Name the system. If the answer is a person, you have costed a manual process as an automated one.
  4. 04Who is accountable when the output is wrong?Name the individual, not the department and not the vendor.
  5. 05Who reviews it, on what sample, how often?Name the person and the percentage. “Spot checks” means nobody.
  6. 06Can the reviewer overrule it, and is the overrule recorded?If overrules are not recorded, you have no way of ever improving the system or of proving it is safe.
  7. 07Which system of record receives the output?Name it. If the output lands in a document, the parallel process has already begun.
  8. 08What number moves, what is it today, and who reads it monthly?The baseline must exist before go-live. After go-live it can no longer be captured.
  9. 09What is the rollback, and who can trigger it without a meeting?A system nobody is authorised to switch off will be tolerated long past the point of usefulness.

Reading the result

7–9 named answersReady to operationalise. The remaining gaps are known and assignable.
4–6 named answersThis is not a technology problem. Close the gaps in the operating model first — building on top of them will only make them more expensive to fix.
0–3 named answersYou have a demonstration, not a project. Either commission the operational design as a piece of work in its own right, or stop.

What good actually looks like

Where to start, in order

  1. Capture the baseline first. Before any tooling. Most AI programmes cannot demonstrate value not because there was none but because nobody wrote down what “before” looked like, and by the time anyone asks, the comparison is gone.
  2. Choose one workflow that already has an owner. Not the most valuable one — the one where a named person is already accountable for the outcome. Ownership is the scarcest input, and inventing it mid-project rarely works.
  3. Design the review step before the automation. How errors are caught determines how much autonomy the system can safely be given, which determines the whole shape of the build. Deciding it last means rebuilding.
  4. Integrate into the system of record before adding a second use case. Breadth before integration is how organisations end up with six pilots and zero operational systems.
  5. Set a stop date at the start. A date on which the programme is either operational or ended. Absent one, pilots do not fail — they persist.

The part that is rarely said out loud

Most pilots should be stopped rather than scaled, and that is a reasonable outcome rather than a failure. The cost of stopping at week six is a licence fee and some staff time. The cost of stopping at month eighteen is a licence fee, an integration, a training programme, a set of processes rebuilt around something that has now gone, and — much the most expensive item — the organisation’s willingness to try again.

The cheap outcome is an early, explicit, documented no. Very few organisations are set up to produce one, because nobody is rewarded for it and the mechanism for saying it does not exist. Building that mechanism is worth more than the next pilot.

Independent research published through 2025 — including MIT’s NANDA initiative on enterprise generative-AI adoption, and Gartner’s forecasting on proof-of-concept abandonment — put the share of enterprise AI pilots reaching measurable business impact strikingly low. Treat the specific figures as directional; the failure mode they describe is the consistent finding.