Ask yourself one question before you read the rest of this. The AI pilot you are sponsoring right now, the one with the steering committee and the eight-month timeline and the readout deck scheduled for Q3, what decision is it actually helping you avoid?

If you cannot answer that in one sentence, the pilot is not a learning instrument. It is a decision deferral mechanism dressed as a learning instrument. That is not an opinion. After 17 years inside large transformations, it is the single most predictive signal I have for which enterprises will still be running the legacy process in 2028.

The MIT Sloan number, 95% of generative AI pilots returning zero P&L impact, is not the bug. It is the system working as designed. Pilots are not failing. They are succeeding at their actual purpose, which is to absorb budget, occupy calendars, produce a defensible artefact, and push the irreversible decision out by another fiscal year. The team that runs the pilot gets promoted on the strength of the deck. The legacy system stays in place. So does the pilot.

This article is about how that system works, why AI specifically breaks it, and the forcing-function alternative.

Where the Pilot Reflex Came From

The pilot model is not stupid. It is just old. It was built for a world where deploying a new technology meant pouring concrete, training thousands of operators, or replacing a system of record. In that world, the cost of being wrong at scale was so high that running a contained experiment first was straightforward risk management. Six Sigma formalised it in manufacturing. ERP rollouts inherited it in the 1990s. The Lean Startup repackaged it for digital products in the 2010s, although most enterprises misread the lesson and kept the pilot but dropped the bias to ship.

Three things were true in those eras that are not true for AI.

First, the technology was relatively static across the pilot window. The ERP system you piloted in January was the same ERP system you deployed in December. You could learn from the pilot.

Second, the cost of a real deployment was high enough that the pilot's cost was a rounding error. The pilot was a cheap insurance policy on a very expensive decision.

Third, the org changes required were largely process changes, which the line could absorb without rebuilding the operating model. The pilot tested process. The deployment tested scale.

None of those three conditions hold for AI.

The Four Hidden Functions of an AI Pilot

If you sit in enough pilot review meetings inside a Fortune 500, a pattern emerges. The conversation almost never centres on the user, the workflow, or the decision being delegated to the model. It centres on something else. Four somethings, actually.

Function one: budget protection. Once a function has secured a pilot budget, the budget itself becomes the asset. Showing the pilot worked protects next year's allocation. Showing the pilot did not work, in many companies, is career-limiting. The pressure on the team is therefore not to discover the truth about the technology. The pressure is to produce a deck that justifies the spend.

Function two: career insurance. Running an AI pilot is currently one of the fastest ways to get visibility on the executive committee. The number of mid-career operators who have rebranded as "AI leads" in the last 24 months is large enough to distort the entire labour market. Their interest is not in the pilot ending. Their interest is in the pilot continuing, expanding, and becoming a permanent function. The MIT Sloan 95% number is, from this angle, a feature.

Function three: vendor evaluation theatre. Many pilots are, in practice, structured procurement exercises. Three vendors are invited. Each runs a contained proof of concept. The output is a vendor selection. The actual deployment, if it happens, often gets re-scoped at the contract stage and bears little relation to the pilot. The pilot served the procurement function. It did not serve the technology decision.

Function four: decision avoidance. This is the deepest one. AI deployment forces uncomfortable questions about which roles change, which processes get retired, which vendors get cut, and which units own which workflows. A pilot lets the executive committee say "we are evaluating" instead of "we are committing." The pilot is the answer to the question the executive committee does not yet want to answer.

I covered the leadership-layer version of this dynamic in Most Chief AI Officers Are Hired to Fail. The pilot is the operational mirror of the ceremonial CAIO. Both are designed to give the appearance of motion without forcing the underlying commitment.

The MIT 95% Number Reframed

When MIT Sloan reported in mid-2025 that 95% of generative AI pilots return zero measurable P&L impact, the trade press read it as a tragedy. I read it as a confession.

If 95% of pilots returned zero P&L impact and the same 95% were still being funded the next year, the pilots were not failing. They were succeeding at something other than what was being measured. They were succeeding at occupying the space where a decision should have been made. They were succeeding at giving the executive committee a defensible answer to the "what are you doing about AI" question for another four quarters. They were succeeding at allowing the operating model to remain unchanged while looking modern.

That is not a failure rate. That is a function.

The companies in the 5% are not better at pilots. They are better at not piloting in the first place. They are deploying. They are accepting the discomfort of the commitment and producing the result that the commitment makes possible. The 5% number is not the survival rate of pilots. It is the rate at which organisations stop pretending.

This is the same dynamic I unpacked from the org-design angle in Why Most Organizations Fail at AI and at the central-function level in The AI Transformation Office Is the Last Job You Should Create. At every layer of the org, the same pattern repeats. Structure absorbs ambition. The ambition never reaches the P&L.

Why AI Specifically Breaks the Pilot Model

Even if you accept that pilots are politically convenient, you might still believe they have learning value. For AI in 2026, they mostly do not. Three reasons.

Capability moves faster than the pilot timeline. A typical enterprise AI pilot runs nine to fifteen months from scoping to readout. METR's 2025 Time Horizon work suggests the underlying capability of frontier models doubles roughly every three months. By the time the pilot reports, the system you piloted is two to four generations behind what is available. The "lessons learned" are largely about a model that no longer exists. The deployment, if it ever happens, will be on a different model with different failure modes.

The model you piloted is not the model you would deploy. Pilots almost always run with a single vendor, a constrained dataset, and a contained user group. Production deployment requires multi-vendor evaluation, full data governance, real user volume, observability, and a fallback plan. Almost none of that is rehearsed in the pilot. The pilot does not de-risk the deployment. It de-risks the wrong thing.

Scope drifts under the team. Every pilot extension widens the scope. By month nine, the team is solving a different problem than the one they scoped in month one. The success criteria have shifted. The original sponsors have rotated. The deck that closes the pilot is, in most cases, a defence of the work done, not a recommendation about what to do next. By the time it lands, it is a museum piece.

The combination of these three is brutal. The pilot does not learn the right things. The pilot does not de-risk the right things. And the pilot's own findings are obsolete by the time they are presented.

The Forcing Function Alternative

If pilots are the disease, what is the medicine? Not bigger pilots. Not better pilots. Forcing functions.

A forcing function is a single decision that closes the off-ramp on every plane of work behind it. It does not ask whether AI will work. It commits the organisation to a date by which the legacy way of working will not exist. Then it lets the operators figure out how to make the AI work, because the alternative is broken process.

Four forcing functions consistently produce real AI deployments inside enterprises.

Production-first deployment. Skip the pilot. Deploy the lowest-risk version of the AI workflow to a real production environment with real users, real volume, and a real fallback. The operators learn what matters in week one because the system has to actually work. There is no decks-only phase. The first deck is the post-implementation review.

Sunset clauses on legacy processes. When you commit to an AI workflow, commit at the same meeting to the date the legacy workflow will be retired. Put it in writing. Communicate it to the line. The forcing function is not the AI. It is the absence of the alternative.

P&L-tied OKRs at the operator level. Tie the operator's quarterly objectives to the unit-economics outcome the AI is supposed to produce. Not "AI capability built." Not "users trained." Cost per transaction, throughput per FTE, contribution margin per region. If the operator's bonus depends on the number, the pilot conversation ends the next morning.

Irreversible commitments. Cancel the legacy vendor contract. Reduce the headcount line in the budget. Sign the new vendor on a multi-year basis. Each of these is a one-way door. They make procrastination structurally impossible. The conversation shifts from "should we" to "how fast."

The Use Case Prioritization tool exists to help with the version of this conversation that happens before the commitment. The AI ROI Calculator is the version that happens at the commitment itself. Neither replaces the executive judgement. Both make the judgement visible.

The Five-Question Diagnostic

Before you fund the next pilot in your organisation, run it through this. If the answer to any of these five is no, what you are funding is not a pilot. It is procrastination with a budget code.

  1. Is there a named, written commitment that if the pilot meets its success criteria, it will be in full production within 90 days, with a sunset date for the legacy process?
  2. Is the operator running the pilot incentivised on the unit economics the pilot is supposed to move, not on "completing the pilot"?
  3. Is the executive sponsor committing to a specific irreversible decision (vendor contract, headcount delta, retired workflow) on the day the pilot reports?
  4. Is the success criterion a single number on the company's books, agreed in advance and not re-negotiable at the readout?
  5. Is the pilot timeline shorter than the underlying technology's capability doubling cycle, so that what you learn is still relevant when you finish?

In my experience inside large organisations, fewer than one in ten pilots passes all five. The rest are funded anyway, because nobody in the room is willing to say the quiet part out loud.

What the Replacement Looks Like in Practice

I will give you the most uncomfortable example I have lived through. A multinational manufacturer wanted to "pilot" an AI-augmented quality inspection workflow. The pilot scope: one line, one product family, six months, three vendors. I argued against it. We did something else instead.

We picked the line. We told the line manager that on a specific date, four months out, the legacy QA process for that line would be retired. We funded the headcount change in the same meeting. We signed a single vendor on a one-year contract with a clear performance clause. The line manager's bonus was tied to defect escape rate and throughput, not to "AI implementation." We did not call it a pilot. We called it a deployment.

In month two, the system was failing on a specific product variant. In month three, the team rebuilt the data pipeline because the original sample was unrepresentative. In month four, the legacy process was retired on the agreed date. By month nine, the workflow was generating measurable margin improvement and had been replicated to four more lines, each with its own sunset commitment.

The entire investment was 30% smaller than the comparable pilot programme another business unit ran in parallel. The pilot programme is still running. It has produced two impressive decks. It has not produced a decision.

This is the difference. One was committed to learning by shipping. The other was committed to learning about whether to commit. They are not the same activity.

The Question I Would Sit With

When I have this conversation with executive committees, the most senior person in the room often pushes back with the same line. "We cannot just commit. We need to learn first."

The honest version of that sentence is different. The honest version is, "I cannot commit, because committing means choosing something to retire, and I am not yet willing to make that choice." That is a defensible position. It is not the same position as "we are piloting."

If you are running an AI pilot today, the question I would sit with is not "is the pilot going well?" The question is the harder one.

What decision is your pilot helping you avoid?

Look at your current AI pilot portfolio. Count the pilots that have been extended once. Count the ones extended twice. Count the ones with no production sunset date, no operator P&L tie, no irreversible commitment in the next quarter.

The number you can honestly count is also the number of decisions your organisation has agreed not to make this year.