The quote is ready. Nobody can sign it.

Imagine a configure-price-quote assistant working exactly as promised. It takes a customer's requirements and returns a configuration, delivery commitment, price and draft response in minutes.

Engineering can test the configuration, finance can inspect the margin and operations can check capacity. The customer, however, receives one combined offer. If the system reached that offer through reasoning nobody can follow before the deadline, who confirms that the whole thing should leave the company?

The usual answer is that an expert will review it. That answer sounds responsible until someone asks which expert, what they are expected to check, how much time they have and whether they can stop the sale.

This is the verification ceiling: the point at which AI produces work faster than qualified people can check the decision it supports. The excess becomes verification debt, meaning output that entered a workflow or reached a customer without confirmation from anyone accountable for it.

I do not expect one person to recall everything a frontier model can retrieve. That is fine. Human oversight does not require the reviewer to know more than the model. It requires a credible way to test the output and the authority to reject it.

Why review is becoming the constraint

The clearest evidence comes from mathematics, which is also where we should be most cautious about generalising to business.

On 1 August, OpenAI reported ten results in mathematics and theoretical computer science. In the company's wording, each result resolves or makes substantial progress on a long-standing open problem. The work spans eight fields. OpenAI says the successful solution tokens would cost roughly $2,000 at Sol API rates.

That is not the cost of the research program. It excludes training, unsuccessful searches and the human work of choosing and formulating the problems. People also prepared the manuscripts and helped formalise the proofs. Each argument came with a Lean certificate, so a machine could check whether the formal proof held.

This is why mathematics is such a useful example. Generation became cheaper, but trust still depended on a separate checking system and people willing to take responsibility for the result. A quotation, delivery promise or commercial recommendation has no equivalent to a Lean proof checker.

METR supplies a second signal. In January, it fitted a doubling time of 88.6 days to the length of tasks that models from 2024 onward complete at 50 percent success. Across the full record, the fitted doubling time was about 196 days.

Those numbers are regression slopes, not a count of observed doublings. METR says the result is sensitive to task composition, and its benchmark mostly contains software and research tasks with clear scoring. It also warns that measurements above 16 hours are unreliable with the current task suite. Fifty percent success is nowhere near an acceptable standard for an industrial process.

I read the result as a warning about operating capacity, not permission to delegate longer tasks. The amount of work a model can attempt is rising quickly. The time of the engineer, commercial manager or service expert who must check it is not rising with it.

The organizers of the NeurIPS 2026 AI for Science workshop describe verification as the new bottleneck. That is the organizers' position, not a consensus research finding. It still captures the management problem well: once generation is abundant, someone must decide which outputs deserve a scarce verification budget.

None of these sources proves that a model can run a commercial workflow unattended. Together, they explain why review capacity deserves a line in the AI business case rather than a vague promise at the end of it.

Put the quote through the Verification Ladder

The useful unit of analysis is the output, not the software. The same assistant can create work that needs very different controls.

The Verification Ladder gives each output a way of being checked, a consequence if wrong and a named signer.

TierWhat the model producedHow it gets checkedConsequence if wrongWho signs
1Draft, summary, readable codeSomeone competent reads itReversible in a dayThe doer
2Code, forecast, configurationHeld-out data, regression suiteReversible in a dayWorkflow owner
3Pricing, routing, targetingControlled release against a baselineRecoverable in a quarterAccountable manager
4Engineering calculation, formal claimFormal checker, or a second independent methodNot recoverableReviewer with no delivery stake
5Novel method, opaque reasoningNot checkable by anyone you haveNot recoverableThe board, explicitly

Checkability sets the minimum tier. Consequence can move it higher. Use whichever produces the stricter control.

Return to the quotation. The configuration may sit at tier 2 if engineering can test it against clear rules. The binding price and delivery commitment belong at tier 3 because an error can take months to recover. A safety or formal engineering claim may push the release to tier 4. The strictest part governs the combined offer.

Tier 5 is not an automatic ban. It is a signal to reduce the system's reach until the output becomes checkable. Shadow mode does not solve the problem because it depends on a human reference decision that tier 5 lacks. Instead, keep the downside survivable, commit in small reversible steps, ask a second team to work from a different starting point, and write the stop condition before anyone sees the model's answer.

The reviewer also needs output they can follow. In 2024, OpenAI researchers used prover-verifier games to study how people check model output under time pressure. When the model was optimized for correctness alone, people became worse at checking it. Training for legibility recovered much of that performance.

That finding belongs in procurement. Ask vendors to show whether a qualified reviewer can follow the output under realistic time pressure. A confident answer and an audit trail are not the same thing.

Independence matters most at tier 4. The people best able to review the work may be the same people who built it or need it approved. If full independence is impossible, record what you have rather than calling every human check independent.

What I would ask before approving the budget

Choose one live workflow where AI already saves time. Do not begin with the whole portfolio. For that workflow, write down five answers:

  1. What exact output reaches a customer or operating decision?
  2. How would a qualified person know that it was wrong?
  3. How long does the check take compared with generation?
  4. Who signs, and can that person stop the release?
  5. What happens when nobody available can verify the output?

These questions turn review from an intention into capacity. If a tier 3 or tier 4 initiative has a model budget but no protected review time, its business case is incomplete. For tier 4, I would also expect a named reviewer with no delivery stake. That is my operating judgment, not an industry benchmark.

Track the share of AI output that reaches a decision without independent confirmation, alongside the review queue, the time it takes to clear and corrections or reversals by verification tier. Those numbers show whether the system is saving time or merely moving work into a queue nobody funded.

Run that review as an operating meeting, not a model meeting. Bring the workflow owner, the person who signs and the person who can change the process. Start with the oldest unchecked output. Decide whether the missing capacity is time, competence, authority or a usable test. Then choose one response: add protected review capacity, make the output easier to check, narrow what the system may do, or stop the release. Record the decision beside the workflow metric. If the same queue returns the following week, treat it as a design defect rather than a temporary workload spike. The point is not to review everything forever. It is to discover which outputs can be tested cheaply, which need a genuinely independent check and which should never have been delegated at their current reach.

The first useful action is small. Take one quotation, configuration or service recommendation from the last week and reconstruct the check. If the team cannot name the method, the signer and the stop rule, do not scale the workflow yet.

The person who signs must be able to stop it

The EU AI Act gives human oversight specific content for high-risk systems.

Article 14 of Regulation (EU) 2024/1689 applies to providers. As appropriate and proportionate, they must make the system overseable by people who can understand its capacities and limitations, monitor anomalies, account for automation bias, interpret the output, override it and intervene or stop the system.

Article 26 applies to deployers. They must assign oversight to people with the necessary competence, training, authority and support.

Put simply, the provider must make oversight possible. The company using the system must provide people who can do the work. Buying the software does not buy that competence or authority for you. Whether a particular system is legally high risk needs its own assessment, but the operating question is useful well beyond regulated cases.

Authority matters because the reviewer is deciding what the company is willing to promise. In the quotation example, the person who can hold the offer has real influence over margin, delivery risk and the customer relationship. Giving someone responsibility without permission to say no is theatre.

Influence is the overlooked asset in this system. AI can move it toward the people who decide which outputs deserve trust. Inside a company, that becomes concrete when a reviewer can stop a price, configuration or claim that nobody can defend.

Before approving the next AI capability, ask to see the verification budget beside the model budget. The plan should name the signer, protect enough of their time to review properly and give them permission to stop the work.

Source and disclosure

The factual claims rely on the primary materials listed below. The NeurIPS workshop page is evidence of the organizers' framing, not a research finding. The quotation is hypothetical. The verification ceiling, verification debt and the Verification Ladder are my framing, not measured findings or a compliance standard.

I used AI to help compare sources, test the argument, draft passages and produce the visuals. I chose the angle, checked the claims, rewrote the article and accept responsibility for the published version. The evidence graphic was checked against the original sources on 24 August 2026. The conceptual visuals are labeled as such.

Written in a personal capacity. Views are my own. This is not legal advice or engineering advice.