In one procurement organization in Basel, a contract-review model did its job well, often better than the median human reviewer. A workflow rule then required a person to sign off every output before it acted, whatever the model's confidence and whatever the contract's value. Volume grew and reviewer attention did not. By month six the reviewers were approving almost everything, and by month seven the flags were going into a queue that functioned as a folder. The audit trail was real, and the audit had stopped.

I have seen versions of this in procurement, claims, contract review, content moderation, KYC and dispatch. My argument is that mandatory human-in-the-loop on every AI workflow has become one of the largest hidden costs of enterprise AI, and that it behaves like technical debt: invisible on the balance sheet, compounding quietly, and more expensive to fix the longer it is left.

The deployment in Basel had not failed. It had succeeded at the wrong objective, producing the appearance of governance instead of value. This article covers where the reflex came from, why the cost compounds, when human review genuinely belongs, and a four-step way to replace blanket review without losing trust.

Where the reflex came from

The human-in-the-loop pattern predates generative AI. It came from three older settings, and each made sense in its own context.

The first was regulated machine learning in finance and healthcare, such as credit decisions, claims adjudication and diagnostic flagging. Regulators required human oversight, the model recommended and the person decided, and the audit trail was the deliverable. It worked because volume was bounded and the models were measurably less accurate than the best practitioners.

The second was early consumer-scale machine learning, such as spam detection, fraud rules and recommendation systems. There the human was usually offline and statistical, reviewing labelled samples and retraining the model. Nobody supervised individual outputs in real time, because that would have destroyed the unit economics within a week.

The third was quality assurance for pre-LLM products such as voice assistants, chatbots and document classifiers. The human loop improved the system over time without acting as a production gate.

When generative AI arrived in 2023, advisory decks merged the three into one prescription: always keep a human in the loop. It sounded responsible, and it was the natural thing to tell a board worried about hallucinations. The default was never engineered, though. Nobody specified what kind of person, in what loop, with what authority, over which decisions, at what volume and with what feedback. The phrase did the work of an architecture without being one.

The over-supervision tax

The curve every executive should be able to draw looks like this.

Supervision intensity runs along the horizontal axis and system productivity up the vertical one. Productivity rises steeply with the first layer of oversight, peaks early and then falls into a long flat tail. Past the peak, every extra review meeting, approval gate and quick check is paid for in three currencies, and all three compound.

The first is throughput. A model that reviews a contract in twelve seconds is no faster than a person if every output waits in a 72-hour approval queue. The workflow's cycle time becomes the reviewer's cycle time, and the gain disappears into the queue.

The second is cognitive load. Harvard Business Review's September 2025 article on what it called "workslop" described AI output that looks finished while quietly pushing work back onto the reviewer, who has to read it, reconstruct the reasoning, decide whether to trust it and then approve or rewrite. That load can exceed the effort of drafting the document from scratch.

The third is skill erosion. Research on what has been called the augmentation trap suggests that when people offload thinking to AI and then over-supervise the result, they lose the expertise the supervision depends on. Over time the "human" in the loop is less often the senior expert who could catch the model's edge cases and more often someone who has only ever seen the model's output.

The tax is more expensive in 2026 because the technology is improving quickly. METR's January 2026 update fitted a doubling time of about three months for the length of tasks frontier models can complete. Capability on hard problems is rising fast, while many organizations spend their AI budget having people re-read AI-generated emails.

The strongest objection: when human review belongs

The strongest objection to this argument is that removing people from AI decisions is reckless, and for some decisions that is correct. Human review is essential for decisions with any of three properties.

The first is irreversibility. Sending a wire transfer, terminating an employee, issuing a public statement or filing a legal pleading cannot be undone, so a human gate is a feature.

The second is external regulation. Lending decisions in banking, treatment decisions in clinical settings and underwriting in insurance often require a named person accountable for the outcome, so human review is a compliance requirement.

The third is asymmetric risk. Where the worst case is qualitatively worse than the best case, especially for the safety of customers or colleagues, human review is a risk control.

That list is short. It does not include drafting an internal email, summarizing a meeting, answering a website FAQ, or flagging a contract clause for a reviewer who has not opened the audit trail in months. When human review is the default for everything, it loses value exactly where it matters, because nobody takes a gate seriously when the gate is everywhere. For consequential outputs, the right design is a named signer with time, a method and authority, which I describe in The Verification Ceiling.

I looked at the same reflex from other angles in Why Most Organizations Fail at AI and The Orchestration Era. The urge to supervise everything is the same urge that produces dozens of chatbots in Teams and no agents on the P&L.

Four anti-patterns

When human review is mandated by default, it tends to collapse into one of four patterns, and none delivers the safety it claims.

Rubber-stamp review happens when output volume exceeds what anyone can meaningfully check, so reviewers approve by default. The audit trail says "reviewed", and nobody read it. The Basel case is the example: a healthy-looking audit trail over a system that was unsupervised in all but name.

Defensive review happens when reviewers read the output but are rewarded for rejecting anything ambiguous, because they are accountable if something goes wrong and invisible if everything goes right. The model's most valuable edge cases, the ones an expert would welcome, are systematically blocked, and the system regresses to the most cautious reviewer's standard.

Ceremonial review exists for the auditor, the regulator or the slide. The reviewer can approve or escalate but cannot edit, redirect or improve the output. The loop is closed in form and open in substance.

Fatigue review starts with a reviewer who cares. After months of near-identical decisions they average a few seconds per item, and the system relies on the weakest form of human attention applied to the largest volume of decisions. It is the most dangerous of the four because everything looks fine until it suddenly is not.

If a deployment shows any of these patterns, the person is attached to the loop without a useful function, and the model might do as well, or better, on its own.

The replacement model

The replacement keeps people involved and changes their position relative to the loop.

Four patterns work consistently, and they combine well.

  1. Human on the loop. The person observes the system for drift, anomalies and edge cases, with authority to halt, retrain or roll it back, and owns its behaviour over time instead of approving individual outputs. Airline cockpits went through a similar shift with fly-by-wire: pilots stopped flying every input and started flying the system that flies the inputs.
  2. Exception-based escalation. The model handles routine cases and classifies its own confidence. Below a defined threshold, the output goes to a person with full context and a recommended decision. Most outputs never need a person, and the ones that do get a focused reviewer with time to do it properly.
  3. Post-hoc audit. A senior reviewer examines a statistically meaningful random sample after the fact, specifically to find model failures. The audit improves the system without gating it.
  4. Adversarial review. A second model with a different architecture or training data reviews the first model's high-stakes outputs. Disagreement escalates to a person, and agreement ships with a logged second opinion.

In all four, the person designs and supervises the gate rather than acting as the gate. The role becomes more senior, with fewer people each covering a much larger surface.

A four-step way to retire blanket review

Follow these four steps in order, without skipping.

  1. Classify decisions rather than workflows. Break each AI deployment down to individual decisions and mark whether each is irreversible, externally regulated or asymmetrically risky. Most will be none of these. The ones that are keep human review on purpose, and the rest become candidates for the replacement patterns.
  2. Instrument before you remove anything. Record what reviewers approve and reject, and why, for about six weeks. That data shows the real decision criteria behind the policy, which are often simpler than the policy and sometimes absent.
  3. Make exception-based escalation the default. Set the confidence threshold deliberately. Outputs above it ship, and outputs below it go to a reviewer with context. Track the escalation volume: if it is high the threshold is too tight, and if it is zero it is too loose.
  4. Add adversarial review for the high-stakes remainder. For decisions that stay irreversible, regulated or asymmetric, use a second model of different lineage to challenge the first, with disagreement escalated to a person who arbitrates. The Change Resistance Review helps identify where the organization is likely to resist the transition before you find out the expensive way.

Done in that order, the supervision tax falls, the system gets safer, and reviewers move into the senior design role they should have had from the start.

Monday move

Pick the AI workflow with the most human approvals per week. For one week, record for each approval how long the reviewer spent, whether anything changed, and whether the decision was irreversible, regulated or asymmetrically risky. Any approval that changed nothing and met none of the three conditions is a candidate for exception-based escalation, and the list becomes the business case for redesigning the loop.