The Verification Ceiling: When AI Produces Faster Than You Can Check

Verification debt means missing required controls, not every output without an individual review. Distinguish sampled work from commitments requiring approval, make checking practical, and keep uncheckable claims from consequential release.

The quote is ready. Nobody can sign it.

Imagine a configure-price-quote assistant working exactly as promised. It takes a customer's requirements and returns a configuration, delivery commitment, price and draft response in minutes.

Engineering can test the configuration, finance can inspect the margin and operations can check capacity. The customer, however, receives one combined offer. If the system reached that offer through reasoning nobody can follow before the deadline, who confirms that the whole thing should leave the company?

The usual answer is that an expert will review it. That answer sounds responsible until someone asks which expert, what they are expected to check, how much time they have and whether they can stop the sale.

This is the verification ceiling: the point at which AI produces work faster than qualified people can check the decision it supports. Verification debt arises when required checks are skipped or remain unresolved while the output is used. A deliberately sampled low-consequence process is different from a binding offer released without its required approval.

I do not expect one person to recall everything a frontier model can retrieve, and oversight does not require it. The reviewer does not need to know more than the model. They need a credible way to test the output and the authority to reject it.

Why review is becoming the constraint

The clearest evidence comes from mathematics, which is also where we should be most cautious about generalising to business.

OpenAI’s August mathematics report describes model-assisted results supported by formal Lean certificates and substantial human work. That gives the checking problem a visible form: producing a plausible argument and independently establishing that it holds are different activities. Commercial commitments rarely have such a complete formal checker.

METR’s task-horizon research provides a separate signal about longer tasks at a stated success threshold. Its mostly software and research benchmark does not establish industrial reliability or the review capacity a company needs. I use it as a reason to investigate the local constraint, not as proof that every AI programme has reached it.

In the quotation, that investigation is straightforward: how many offers arrive, which checks do they require, and how long do qualified people need to clear them? If generation grows while the required checks stay slow, the business case must fund the difference or change the work.

Put the quote through the Verification Ladder

The useful unit of analysis is the output rather than the software. The same assistant can create work that needs very different controls.

The following ladder is my routing aid. Its levels are illustrative; they are not validated risk thresholds or an engineering standard.

Output and consequenceAppropriate checking approachAccountable decision
Low-consequence draft, with easy recoveryCompetent review or a documented sampling plan suited to the useThe person using it follows the agreed review rule
Standardised result with known testsApplicable rules, regression tests and explicit exception handlingThe workflow owner maintains the tests and release conditions
Binding price, delivery or other external commitmentThe relevant commercial, technical and capacity checks on the actual caseNamed authority approves the combined commitment
High-consequence technical or formal claimA qualified independent method or reviewer, matched to the claimTechnical approval and release authority are explicit
Output for which no credible check is availableReduce reach; find a valid test or keep the output from consequential useA risk owner may approve a bounded experiment, not certify the answer

Consequence can require stronger controls than ease of checking suggests. An easy-to-read statement can still make an expensive promise. Sampling is appropriate only where the consequence and permitted use justify it; a low-looking tier does not grant permission by itself.

Return to the quotation. Configuration rules may be testable automatically. The delivery date depends on current capacity, and the price depends on the agreement and authorised exceptions. A safety claim may need specialist approval. The combined offer can leave only when its required checks and approvals are complete.

If nobody can verify a consequential part, adding a board signature does not make it correct. The board or another risk owner can decide whether to fund a limited investigation with survivable downside. Release still needs evidence appropriate to the claim. A second team working from a different starting point may help establish that evidence; agreement between two unsupported outputs is insufficient.

The reviewer also needs output they can follow. In 2024, OpenAI researchers used prover-verifier games to study how people check model output under time pressure. When the model was optimized for correctness alone, people became worse at checking it. Training for legibility recovered much of that performance.

That finding belongs in procurement. Ask vendors to show whether a qualified reviewer can follow the output under realistic time pressure. A confident answer with an audit trail can still be impossible to check in the time available.

Independence matters most for high-consequence claims. The people best able to review the work may be the same people who built it or need it approved. If full independence is impossible, record what you have rather than calling every human check independent.

The person who signs must be able to stop it

A named reviewer needs the competence, time, evidence access and authority to do the work. Buying the software supplies none of those automatically. The required controls should be explicit before the team relies on the output.

Authority matters because the reviewer is deciding what the company is willing to promise. In the quotation example, the person who can hold the offer has real influence over margin, delivery risk and the customer relationship. A reviewer who cannot say no is only adding a signature.

AI can shift influence toward the people who decide which outputs deserve trust. Inside a company, that becomes concrete when a reviewer can stop a price, configuration or claim that nobody can defend.

Give the review queue an owner

Take the oldest unresolved quotation to an operating review with the process owner, the proposed signer and the person able to change the handoff. Determine whether it is missing information, a usable test, expert time or a decision right. Those causes need different responses.

The response might be better source records, a testable configuration output, protected review time or a narrower commitment. It might also be stopping release. Record the decision beside the elapsed time and rework for the complete case so that the faster draft does not hide a slower customer answer.

Measure the controls actually required for each class of output. For commitments requiring individual approval, count releases missing that approval and unresolved exceptions carried forward. For deliberately sampled work, track sample coverage, detected failures and the response to them. Report those separately; they cannot be collapsed into one universal percentage of unreviewed output.

Watch arrival rate, clearance time, corrections and review cost together. A recurring queue deserves investigation, not an automatic diagnosis after two weeks. Workload peaks, absence, data failures and poor workflow design can produce similar symptoms. Establish which constraint is recurring before deciding to add people or redesign the process.

For the hypothetical quotation, useful evidence would show that engineering’s checks apply to the correct revision, finance has approved the commercial exceptions, and operations has accepted the delivery promise. It should also show that the signer received those results in time and could hold the offer. A log containing hundreds of tool calls is not a substitute for that handover.

When assessing vendors, ask a qualified reviewer to check an output under the time pressure they would face in the real process. Watch what evidence they need and what they cannot establish. That is a better procurement test than asking whether the answer looks professional.

Review can be lightweight where consequences and recovery permit. The aim is to protect scarce expertise for the decisions that need it. Before expanding the next AI workflow, put its checking method, available capacity and release authority beside its generation cost. The customer receives the combined promise, and somebody has to be able to defend it.

Source and disclosure

This is an editorial synthesis, not an original study of review capacity. The quotation is hypothetical and the ladder is a proposed routing aid. The prover-verifier study concerns legibility in its tested setting; it does not validate these operating categories. The editorial photograph is AI-generated.

Explore topics

Evaluate my AI transformation leadership

For companies assessing an AI transformation leadership appointment: review the public operating case, the proposed mandate and the responsibilities your enterprise needs.

Key takeaways

  • Measure unresolved or missing required checks by output class.
  • Sampling and individual approval require different measures.
  • Risk acceptance does not certify an uncheckable answer.
  • A reviewer needs a usable method, time and authority to hold release.
  • Investigate recurring queues before diagnosing their cause.

Decision brief

Fund and test the required controls before increasing the flow of consequential outputs.

Decision criteria

  • The method can establish what the actual claim requires.
  • Review capacity can meet the release conditions.
  • Someone can hold or reject the output.

Evidence to check

  • Missing approvals and unresolved exceptions by output class.
  • Sample coverage and detected failures.
  • Review arrival rates, clearance time and corrections.

What I would do Monday

Reconstruct one actual release with its required checks, evidence, decision authority and time consumed.

Questions executives ask

Does every AI output need individual approval?
No. A documented sampling approach may be appropriate for some low-consequence uses. Binding commitments and high-consequence claims need the checks defined for their actual use.
Who approves something nobody can verify?
A risk owner can authorise a bounded investigation with controlled downside. That authorisation does not establish correctness or permit an unsupported consequential release.
How do we measure verification debt?
Count missing required checks and unresolved exceptions carried into use, separately for each output class. Report sampling coverage and detected failures separately.

Canonical URL: https://juanbeltran.ch/blog/ai-verification-ceiling