Mandatory human-in-the-loop on every AI workflow has quietly become the single biggest tax on enterprise AI value, and almost nobody is naming it that way.

That is not a careful claim. It is a blunt one, and I will defend it. After 17 years inside large operating models I have watched a specific pattern repeat across procurement, claims, contract review, content moderation, KYC, and dispatch. The model is good. The model is, in fact, often better than the median human reviewer. Then a workflow rule gets bolted on that says every output, regardless of confidence, regardless of materiality, must be signed off by a person before it acts. Volume scales. Reviewer attention does not. By month six the humans are rubber-stamping. By month nine they are on stress leave or have quietly moved to "spot-check" mode, which is the polite way of saying nobody is reading anything. The flags go into a queue that, in practice, is a folder. The audit trail is real. The audit is not.

The deployment has not failed. It has succeeded at the wrong objective. It has shipped governance theatre instead of value. And like technical debt, the cost compounds in the dark. This article is about why that pattern emerged, why HITL has the same compounding properties as technical debt, and the four-step move to retire it without losing trust.

Where the HITL Reflex Came From

The "human in the loop" pattern was not invented for generative AI. It came from three older worlds, and each one made sense in its original context.

The first world was regulated machine learning in finance and healthcare. Credit decisioning, claims adjudication, diagnostic flagging. There, the regulator did not just permit human oversight, they required it. The model was a recommender, the human was the decider, and the audit trail was the deliverable. That pattern worked because the volume was bounded, the decisions were already adjudicated by humans, and the model was demonstrably less accurate than the best practitioners.

The second world was early production machine learning at consumer scale. Spam detection, fraud rules, recommendation systems. There, the human in the loop was usually offline and statistical. They reviewed labelled samples, retrained the model, and shipped a new version. Nobody supervised individual outputs in real time. That would have collapsed the unit economics in a week.

The third world was UX testing in pre-LLM AI. Voice assistants, chatbots, document classifiers. The human-in-the-loop was a quality assurance loop, not a production gate. It improved the system over time. It did not bottleneck it.

When generative AI arrived in 2023, every consulting deck collapsed those three traditions into one prescription. "Always keep a human in the loop." It sounded responsible. It sounded safe. It sounded like the natural thing to say to a board that was nervous about hallucinations and reputational risk. So it became the default.

The problem is that the default was never engineered. Nobody asked, "what kind of human, in what kind of loop, with what kind of authority, on what kind of decisions, at what volume, with what kind of feedback signal?" The phrase did the work of an architecture without being one.

The Over-Supervision Tax

Here is the curve every executive needs to be able to draw.

On the X axis, supervision intensity. On the Y axis, system productivity. The curve rises steeply with the first bit of oversight, peaks early, and then collapses into a long flat tail. Past the peak, every extra review meeting, every extra approval gate, every extra "let me just check this" is a tax on the gain. It is not free. It is paid in three currencies, and all three compound.

The first currency is throughput. A model that can review a contract in 12 seconds is not faster than a human if every output sits in a 72-hour approval queue. The cycle time of the workflow becomes the cycle time of the human, and the productivity gain disappears into the queue.

The second currency is cognitive load, and this is the one Harvard Business Review named in its September 2025 piece on what it called "workslop." AI-generated outputs that look finished but secretly transfer cognitive load back to the reviewer. The reviewer now has to read, reconstruct the model's reasoning, decide whether they trust it, and either rubber-stamp or rewrite. The honest version of that load is often higher than if they had drafted the document themselves.

The third currency is skill atrophy. The MIT Augmentation Trap research, work by Caosun and Aral and others, shows that when humans cognitively offload to AI but then over-supervise the result, they erode the very expertise that the supervision depends on. Two years in, the team is slower than before, less skilled than before, and more dependent on the model than before. The "human" in "human in the loop" is no longer the senior expert who could catch the model's edge cases. It is a junior who only ever saw the model's output and now cannot remember what good looked like before.

What makes this tax especially expensive in 2026 is the second half of the picture. METR's January 2026 Time Horizon report shows AI task complexity doubling roughly every three months. Frontier systems are now solving problems that humans struggled with for years. Hold both facts at the same time. The technology is getting demonstrably better at hard problems faster than any technology in living memory. And most organisations are spending their AI budget making humans re-read AI-generated emails.

That is what I mean when I say HITL is the new technical debt. It is not visible on the balance sheet. It accumulates quietly. It compounds. And every quarter you do not refactor it, the cost of refactoring it goes up.

When HITL Still Belongs

I want to be precise about this, because the reflex when somebody attacks HITL is to assume they want a free-running model with no oversight. That is not the argument.

There is a small, important set of decisions where HITL is not just acceptable but mandatory. They share three properties.

They are genuinely irreversible. Sending a wire transfer. Terminating an employee. Issuing a public statement. Filing a legal pleading. Once the action is taken, there is no rollback. For these, HITL is a feature.

They are regulated by an external authority. Lending decisions in banking. Treatment decisions in clinical settings. Underwriting in insurance. Most jurisdictions explicitly require a named human accountable for the outcome. For these, HITL is a compliance requirement.

They carry asymmetric reputational or safety risk. Anything that could put a customer at risk. Anything that could put a colleague at risk. Anything where the worst-case outcome is qualitatively worse than the best-case outcome. For these, HITL is a risk control.

That is the list. It is not long. It does not include "drafts an internal email to my team." It does not include "summarises a meeting." It does not include "answers an FAQ from a website visitor." It does not include "flags an unusual contract clause for a reviewer who has not opened the audit trail in six months." When you make HITL the default for everything, you devalue it in the places where it actually matters. Nobody takes the gate seriously when the gate is everywhere.

I covered the same fundamental dynamic, framed differently, in Why Most Organizations Fail at AI and in The Orchestration Era. The reflex to supervise everything is the same reflex that produces forty chatbots in Teams and zero agents on the P&L.

The Four Anti-Patterns

In practice, when HITL is mandated as the default, it collapses into one of four anti-patterns. I have seen all four many times. None of them produce the safety they claim.

Rubber-stamp review. The volume of model outputs exceeds what any human can meaningfully review, so reviewers approve everything by default. The signature on the audit trail says "reviewed by." The reality is that nobody read it. This is the procurement contract example from Basel. The audit trail looks healthy. The system is unsupervised in everything but name.

Defensive review. The reviewer reads the output but their incentive is to reject anything ambiguous. Why? Because they are accountable if something goes wrong and irrelevant if everything is fine. Asymmetric incentives produce asymmetric behaviour. The result is that the model's high-value edge cases, the ones a human expert would have welcomed, are systematically blocked. The system regresses to the mean of the most cautious reviewer.

Theatre review. The review exists for the audit, the regulator, or the slide. The reviewer reads the output, but the workflow gives them no real authority to alter it. They can approve or escalate. They cannot edit, redirect, or improve. The "loop" is closed in form and open in substance. The customer never sees the difference.

Fatigue review. The reviewer started caring. After six months of identical decisions, they stopped. After twelve, they were averaging ten seconds per item. After eighteen, they were on sick leave. The system is now relying on the worst version of human attention, applied to the highest volume of decisions. This is the most dangerous anti-pattern of the four because it looks like everything is working until it suddenly is not.

If your AI deployment has any of these four patterns, you do not have a human in the loop. You have a human glued to the loop with no useful function. The model would have done better, and possibly safer, alone.

The Replacement Model

The replacement is not "remove the human." That is the strawman the reflex defenders of HITL always reach for. The replacement is to put the human in a fundamentally different position relative to the loop.

There are four patterns I have watched work consistently. They are not exclusive. They compose.

Human-on-the-loop. The human is not in every decision. They are observing the system, watching for drift, anomalies, and edge cases. They have authority to halt, retrain, or roll back. They do not approve individual outputs. They own the system's behaviour over time. This is the same shift that happened in airline cockpits when fly-by-wire arrived. Pilots stopped flying every input. They started flying the system that flew the inputs. The result was safer, not less safe.

Exception-based escalation. The model handles the routine. The model itself classifies its own confidence. Below a defined threshold, the output is escalated to a human, with full context and a recommended decision. Most outputs never see a human. The few that do, see a focused human with the time to review properly. The throughput stays high. The attention stays sharp.

Post-hoc audit. A statistically significant random sample of outputs is reviewed after the fact, by a senior reviewer, with the explicit purpose of finding model failures. The audit improves the system. It does not gate the system. The customer is served at machine speed. The improvement happens on a slower clock.

Adversarial AI review. A second model, with a different architecture or training set, reviews the first model's high-stakes outputs. The two disagreeing is the signal that escalates to a human. The two agreeing is the signal that ships. This is the pattern frontier labs are now using on their own model outputs. It scales in a way that human review cannot.

In all four patterns, the human is the architect of the gate, not the gate itself. The role is more senior, not less. The compensation is higher, not lower. There are fewer of them, but each one has a much larger surface of impact. This is what good organisational design has always looked like, applied to a new technology.

A Four-Step Move to Retire HITL Safely

If you want to leave a Friday like the one in Basel behind, here is the sequence I recommend, in this order, with no skipping.

Step one: classify your decisions, not your workflows. Take each of your AI deployments and break it down to the level of individual decisions. For each decision, mark whether it is genuinely irreversible, externally regulated, or asymmetrically risky. Most will not be. The ones that are stay HITL, on purpose. Everything else is now a candidate for one of the four replacement patterns.

Step two: instrument before you remove. Before you take a human out of any loop, instrument the loop. Record what the human was approving, what they were rejecting, and on what grounds. Six weeks of that data tells you the actual decision-making criteria, separate from the policy claims. In most cases, the criteria are far simpler than the policy. Sometimes they are non-existent. Both findings are useful.

Step three: shift to exception-based escalation as the default. Once you have instrumented, set the model's confidence threshold deliberately, not by feel. Outputs above the threshold ship. Outputs below escalate, with context, to a human who has the time to review them properly. Track the volume of escalations. If it is high, the threshold is too tight. If it is zero, the threshold is too loose. This is a control surface. Use it.

Step four: build adversarial review for the high-stakes residue. For the small number of decisions that stay irreversible, regulated, or asymmetric, use a second model with a different lineage to challenge the first. Disagreement escalates. Agreement ships, with a logged second-opinion trail. The human becomes the arbitrator of disagreement, not the reviewer of every output. This is more rigorous than HITL, not less. It is also dramatically more scalable. The Change Resistance Review helps you examine where the organisation may fight this transition before you find out the expensive way.

If you do those four steps in that order, the supervision tax collapses, the system gets safer, and the human reviewers move into the senior architectural role they should have been in from the start.

The Real Question on the Table

I keep coming back to that procurement director in Basel. Her organisation had not failed because of a bad model. The model was excellent. They had failed because the operating model surrounding the model treated the human reviewer as a safety feature when, by month seven, the human reviewer was the source of the risk.

The frontier models are not slowing down. The over-supervision tax is not getting cheaper to pay. And the organisations that retire HITL deliberately, replace it with something more rigorous, and re-architect the human role around exception-handling, audit, and adversarial review will pull away from those that keep the default in place out of habit.

So here is the question I would put to the next executive team that asks me how to scale AI safely.

For each step in your AI workflow where a human is currently approving an output, are they actually doing work, or are they only absorbing accountability?

The honest answer to that question is also the answer to whether your organisation will be in the 5% or the 95%.