Think of the AI pilot you are sponsoring right now, the one with the steering committee, the eight-month timeline and the readout deck scheduled for Q3. What decision is it helping you avoid?

If you cannot answer that in one sentence, the pilot is probably deferring a decision rather than informing one. In the transformations I have worked on, that is the most reliable signal I know for which enterprises will still be running the legacy process in 2028.

MIT NANDA's GenAI Divide report found that about 95% of generative AI pilots return no measurable P&L impact. The usual reading is that pilots are failing. My reading is that many of them succeed at a different purpose: they absorb budget, occupy calendars, produce a defensible artefact and push the irreversible decision out by another fiscal year. The team that runs the pilot is rewarded for the deck, and the legacy system stays in place.

This article is about how that system works, why AI breaks it, and what to use instead.

Where the pilot reflex came from

The pilot model made sense for the world it was built for. Deploying a new technology used to mean pouring concrete, training thousands of operators or replacing a system of record, so the cost of being wrong at scale justified a contained experiment first. Six Sigma formalized that in manufacturing, ERP rollouts inherited it in the 1990s, and the Lean Startup repackaged it for digital products in the 2010s, although many enterprises kept the pilot and dropped the bias to ship.

Three things were true in those eras that are not true for AI. The technology stayed roughly the same during the pilot, so the ERP system piloted in January was the one deployed in December. The real deployment was expensive enough that the pilot was cheap insurance. And the organizational changes were mostly process changes the line could absorb without rebuilding the operating model.

Four functions a pilot can serve

Sit in enough pilot reviews and a pattern emerges. The conversation rarely centres on the user, the workflow or the decision being delegated to the model. It centres on four other things.

The first is budget protection. Once a function has secured pilot money, the budget becomes the asset. Showing that the pilot worked protects next year's allocation, and showing that it did not can hurt careers, so the pressure is to justify the spend rather than to learn the truth about the technology.

The second is career visibility. Running an AI pilot is one of the fastest routes to the executive committee's attention, and many mid-career operators have rebranded as AI leads. Their incentive is for the pilot to continue, expand and become a permanent function.

The third is procurement. Many pilots are structured vendor evaluations: three vendors run contained proofs of concept and the output is a selection. The deployment that follows, if any, is often re-scoped at contract stage and bears little relation to the pilot.

The fourth is decision avoidance, and it runs deepest. AI deployment forces uncomfortable questions about which roles change, which processes are retired, which vendors are cut and which units own which workflows. A pilot lets the executive committee say "we are evaluating" instead of "we are committing".

I covered the leadership version of this dynamic in Most Chief AI Officers Are Hired to Fail. A pilot without a decision gate is the operational mirror of a ceremonial Chief AI Officer: both give the appearance of motion without forcing the commitment.

Reading the 95% differently

The trade press read the MIT NANDA finding as a tragedy.

If most pilots return no P&L impact and are funded again the following year, they are delivering something other than what is being measured. They fill the space where a decision should have been made, give the executive committee an answer to "what are you doing about AI" for another four quarters, and let the operating model stay the same while looking modern.

The companies in the successful minority are not better at pilots. More often they deploy, accept the discomfort of the commitment and get the result the commitment makes possible.

I looked at the same dynamic from the organization-design angle in Why Most Organizations Fail at AI and at the central-function level in The AI Transformation Office Is the Last Job You Should Create. At each layer, structure absorbs ambition before it reaches the P&L.

Why AI breaks the pilot model

Even if pilots are politically convenient, they might still have learning value. For AI in 2026 they often do not, for three reasons.

Capability moves faster than the pilot. A typical enterprise AI pilot runs nine to fifteen months from scoping to readout, while METR's January 2026 update fitted a doubling time of about three months for the length of tasks frontier models can complete. By the time the pilot reports, the system it tested is likely to be several model generations old, and the deployment would run on a different model with different failure modes.

The pilot tests the wrong system. Pilots usually run with one vendor, a constrained dataset and a small user group. Production needs multi-vendor evaluation, full data governance, real volume, observability and a fallback plan, and almost none of that is rehearsed. The pilot reduces risk in the wrong place.

Scope drifts. Every extension widens the scope, and by month nine the team is solving a different problem from the one scoped in month one. The success criteria have moved, the original sponsors have rotated, and the closing deck defends the work done rather than recommending what to do next.

Forcing functions instead

The alternative to a bigger or better pilot is a forcing function: a decision that closes the off-ramp for the work behind it. It commits the organization to a date after which the legacy way of working will not exist, and then lets operators work out how to make the AI work, because the alternative is a broken process.

Four forcing functions consistently produce real deployments.

  1. Production-first deployment. Deploy the lowest-risk version of the AI workflow into production, with real users, real volume and a real fallback. Operators learn what matters in the first week because the system has to work, and the first deck is the post-implementation review.
  2. Sunset clauses. When you commit to an AI workflow, set the retirement date of the legacy workflow in the same meeting, put it in writing and tell the line.
  3. Operator OKRs tied to the P&L. Tie the operator's quarterly objectives to the unit economics the AI is meant to change, such as cost per transaction, throughput per FTE or contribution margin per region, rather than to "capability built" or "users trained".
  4. Irreversible commitments. Cancel the legacy vendor contract, reduce the headcount line or sign the new vendor for a multi-year term. Each is a one-way door that shifts the conversation from whether to how fast.

The Use Case Prioritization tool supports the conversation before the commitment, and the AI ROI Calculator supports the commitment itself. Neither replaces executive judgment, but both make it visible.

The five-question diagnostic

Before funding the next pilot, run it through these five questions. If any answer is no, you are probably funding delay.

  1. Is there a named, written commitment that the work goes into full production within 90 days if it meets its success criteria, with a sunset date for the legacy process?
  2. Is the operator running it incentivized on the unit economics it is supposed to move, rather than on completing the pilot?
  3. Will the executive sponsor make a specific irreversible decision, such as a vendor contract, a headcount change or a retired workflow, on the day it reports?
  4. Is the success criterion a single number on the company's books, agreed in advance and closed to renegotiation at the readout?
  5. Is the timeline shorter than the technology's capability cycle, so that what you learn is still relevant when you finish?

In my experience, fewer than one pilot in ten passes all five, and most are funded anyway because nobody in the room wants to say so.

What the replacement looked like

The clearest example I have lived through involved a multinational manufacturer that wanted to pilot AI-assisted quality inspection: one line, one product family, six months and three vendors. I argued against it, and we did something else.

We picked the line and told the line manager that the legacy QA process for that line would be retired on a specific date four months out. We funded the headcount change in the same meeting, signed one vendor on a one-year contract with a clear performance clause, and tied the line manager's bonus to defect escape rate and throughput rather than to "AI implementation". We called it a deployment.

In month two the system failed on one product variant. In month three the team rebuilt the data pipeline because the original sample was unrepresentative. In month four the legacy process was retired on the agreed date. By month nine the workflow was producing measurable margin improvement and had been replicated on four more lines, each with its own sunset commitment. The investment was about 30% smaller than a comparable pilot programme another business unit ran in parallel, which is still running, has produced two impressive decks and has not produced a decision.

The strongest objection

A risk-conscious executive will say that pilots exist for good reasons. Putting untested AI into production is reckless in regulated or safety-critical work, auditors and regulators want evidence before deployment, and a forcing function can force a bad deployment as easily as a good one.

That is right for high-consequence workflows. Where an output commits a price, a safety claim or a regulated decision, the verification discipline I describe in The Verification Ceiling has to come first, and production-first should mean the lowest-risk slice with a fallback. My objection is to pilots with no decision gate, no owner incentive and no retirement date. A controlled experiment with those three things is a legitimate step toward a decision.

Monday move

Look at your current AI pilot portfolio and count three things: the pilots that have been extended at least once, the ones with no production or sunset date, and the ones where no executive has committed to an irreversible decision in the next quarter. Take the oldest pilot on that list to your next leadership meeting and ask for one of two decisions: set the date it goes into production with a sunset for the legacy process, or stop it.