I do not know more than the model
Nobody says this out loud in a steering committee.
On recall and breadth of retrievable knowledge, I do not beat the model. Neither do you. Neither does your best engineer or your most experienced quality manager. That is a claim about retrieval, not about judgment, and it holds hardest in the domains where an answer can be checked.
That contest is over. It ended quietly, without an announcement, and most organizations are still hiring, promoting, and organizing as though it were still running.
If knowledge is no longer the scarce thing, something else is.
The scarce thing is verification. Not the ability to produce an answer, but the ability to tell whether an answer is right.
I argued the operational half of this in Human-in-the-Loop Is the New Technical Debt: review does not scale. This is what happens after that stops being a capacity problem. A throughput bottleneck is solved by adding reviewers. A ceiling is not, because output can now exceed the best available reviewer's capability rather than merely their capacity. Different problems, different money.
Tool is the wrong word
We call this a tool because tool is the word we have for new technology at work. It is the wrong word, and wrong words produce wrong budgets.
The internet moved information from one place to another. That was the entire miracle, and it was enormous. Transport.
This does something else. It turns the same information into something that was not there before. A complaint becomes a root cause. A constraint becomes an option. Transformation.
Transport requires no judgment: a file arrived or it did not. Transformation requires someone to say whether it was correct.
So ask the operational version about your own company. If every model your teams touch went dark on Monday morning, would the work be slower, or would some of it simply stop? In a growing number of industrial B2B functions I expect the answer is increasingly the second one. Technical documentation. First-line service triage. Code. Translation. Market scanning.
That is not a tool. That is a dependency.
You do not write a use case for a dependency. You write a verification requirement.
The verification ceiling: where the two curves cross
For a while there was a comfortable story, and I told a version of it myself: these systems are bounded by their corpus, they recombine rather than originate, and whatever they produce a sufficiently expert human could have produced.
On 1 August, OpenAI reported ten results in mathematics and theoretical computer science, each of which, in its words, resolves or makes substantial progress on a long-standing open problem. They span high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. Three are Erdos problems. OpenAI puts the tokens needed to find those solutions at roughly $2,000 at its list API rates, which is a legibility figure rather than a cost anyone actually paid: it excludes training, the searches that found nothing, and the human work of choosing and formulating the problems.
At the International Mathematical Olympiad in Shanghai in July, Huawei and Xiaohongshu each announced that their models had solved all six problems for 42 out of 42. Both companies say the answers were submitted to the organizers for official grading. Reporting on the event puts the number of human contestants at full marks at seven of 666, and notes that several other frontier models are reported to have matched the score unofficially on the same problems. These are company claims relayed by news organizations, not primary disclosure, and I have found no jury statement confirming them.
Track the trajectory rather than the headline. In 2024, two dedicated research systems reached 28 of 42, silver-medal standard, and it was treated as a landmark. In 2025, 35 of 42. This year, reportedly full marks.
Mathematics is unusually friendly to machine search precisely because correctness can be formally checked, and none of it generalizes automatically to your pricing model.
But the specific claim that the training corpus is the ceiling is finished. It has been for a while, and most operating models have not caught up.
Here is what that does to the shape of the problem.
If a system can produce a result no human produced, then "is this correct?" can no longer be answered by asking someone senior. It has to be answered by checking. Checking is a human activity that runs at human speed.
That is the shape I keep seeing in stalled AI programs, and I want to be clear that this is my read of a pattern rather than a survey. The model produces at machine speed. The review queue runs at human speed. The queue becomes the roadmap. The failure mode from there is predictable: the queue gets treated as the problem, and review is loosened rather than resourced. That is how a governance failure gets filed as a productivity win.
Only one side of this has a growth curve at all. In January, METR fitted a doubling time of 88.6 days to models from 2024 onward for the length of task an agent completes at 50 percent success, against 196 days across the full record. Both are regression slopes rather than counted doublings, and METR notes part of the change may reflect its new task mix. Fifty percent is a benchmark threshold, not an operating standard: at the reliability an industrial process accepts, the horizon is far shorter. The claim is not that agents can be trusted with long tasks. It is that the length of work being handed over is growing on a curve and the review function is not. My reading of the two windows is that the curve is bending, which is an inference from a fit rather than a measurement.
A smaller detail in the same work is worth more than the headline. METR notes that measurements above 16 hours are unreliable with its current task suite. Our ability to measure the capability and our ability to check its output are failing for the same reason.
The verification ceiling is the point at which an organization can no longer check what its AI systems produce as fast as they produce it. Above that line, output accumulates as verification debt: work that was generated, was probably useful, and was never independently confirmed by anyone accountable for it.
The organizers of the NeurIPS 2026 AI for Science workshop, in their framing statement, write that the bottleneck for AI for Science is no longer hypothesis generation but verification, and set the central problem as which AI outputs deserve a scarce verification budget. That is seven researchers scoping a workshop, not a finding. I read it as a signal that the problem has a name now.
Most companies I meet are still funding capability and starving verification, then describing the result as an adoption problem.
The Verification Ladder: tiering every AI output
There will be answers I cannot check. Not difficult to check. Cannot. A method whose reasoning does not decompose into steps I can follow, arriving at a conclusion that is either a significant advantage or a significant liability, and I will not be able to tell which by looking at it.
Both obvious moves are bad. Refuse everything you cannot verify, and you forfeit the gains to whoever has a weaker conscience. Approve what you cannot verify, and you have not delegated a task; you have transferred authorship of the decision to something that cannot be held responsible for it.
So treat verification as capacity you build before you need it, not a step at the end.
One clarification, because "verification" is doing too much work in most conversations. Your organization already runs four activities with different cost curves and owners: verification that a thing was built to spec, validation that it is fit for its intended use, independent assurance for a regulator or customer, and peer review. Tiers 1 and 2 below are verification, tier 3 is validation, and tiers 4 and 5 are independent assurance, which is the expensive one because it ships no feature.
The Verification Ladder assigns every AI output a level of checking and a named signer, from checkable by inspection at tier one to checkable by nobody available at tier five.
| Tier | What the model produced | How it gets checked | Consequence if wrong | Who signs |
|---|---|---|---|---|
| 1 | Draft, summary, readable code | Someone competent reads it | Reversible in a day | The doer |
| 2 | Code, forecast, configuration | Held-out data, regression suite | Reversible in a day | Workflow owner |
| 3 | Pricing, routing, targeting | Controlled release against a baseline | Recoverable in a quarter | Accountable manager |
| 4 | Engineering calculation, formal claim | Formal checker, or a second independent method | Not recoverable | Reviewer with no delivery stake |
| 5 | Novel method, opaque reasoning | Not checkable by anyone you have | Not recoverable | The board, explicitly |
Checkability sets the floor. Consequence sets the ceiling. Take the higher of the two.
That rule matters because output type alone will not classify real work. Our own configure-price-quote assistants emit a configuration, which is tier 2 by type, and a binding price, which is tier 3, and the higher one governs. A remaining-useful-life model emits a forecast, tier 2 by type, carrying a safety consequence that puts it at tier 4.
Tier five is not a ban. It is an instruction: reduce reach until the output drops a tier. Shadow mode is the wrong tool here, because it needs a human reference decision that tier five by definition lacks. What works without a baseline is bounded downside so being wrong is survivable, staged commitment into the smallest reversible increment, adversarial replication by a second team starting from a different place with disagreement as the escalation trigger, and a stop threshold written before the output exists. Write that threshold in advance or the board tier becomes a rubber stamp with better minutes.
The limit of this argument is worth stating. Formal verification works beautifully in mathematics and in parts of engineering. Most commercial decisions have no proof checker and are not going to get one. A quotation, a capacity commitment, a safety claim, a price: none can be discharged to a certificate. Where no formal checker exists, tier 4 is discharged by a second independent method and a reviewer with no delivery stake, and that is the version most of your work will use.
One practical problem the org chart does not solve: the only people qualified to check a tier 4 output are usually the people who built it. Your options are external review, rotation across sites, or explicitly accepting graded independence and recording which grade you got. Functional safety already grades this as independent person, department, organization. Borrow it, because a graded scale is budgetable and a binary one is aspirational.
What I would actually do on Monday
Take the three AI initiatives your organization is proudest of. For each one, answer five questions in writing.
- What does this produce, and how would we know if it were wrong?
- Who checks it today, how long do they have, and what would they need to already know?
- What is its tier on the ladder, honestly, on the worst day rather than the demo day?
- Are we the provider or the deployer of this system? The answer is per system, not per company, and it decides which obligations are yours.
- What percentage of this initiative's budget is spent producing output, and what percentage confirming it? If you cannot answer in one minute, the answer is close to zero.
My own rough rule, offered as judgment and not as a benchmark: anything at tier 3 or above should carry a review line in its business case, and I would not fund a tier 4 initiative whose plan does not name an independent reviewer with protected time. Put a number on it before somebody else puts a zero there.
Then convert this from a document into a decision. Two of the five answers will be uncomfortable. Bring those two to the next steering committee with a number attached, and name the one initiative you will stop this quarter because nobody can check it.
If question four or five has no answer, you do not have a governance gap. You have an unowned decision, and it will be made under deadline pressure by whoever happens to be closest to it.
The August split: where verification scaled and where it did not
Those ten results contain both halves of this argument.
OpenAI did not only publish claims. It states that the model formalized each argument in a Lean certificate, released alongside the manuscripts, so the correctness of the mathematics can be checked by a machine rather than taken on trust.
Where a formal checker exists, the correctness check itself scales at machine speed. OpenAI is explicit that people were still in the loop: it says it helped prepare the manuscripts and formalize the proofs in Lean, and that it takes responsibility for their correctness. So even the clean case has a human step. It is a far smaller one than reading a 249-page argument, which is the point.
Now the other half. Formal correctness is not understanding. A certificate tells you the argument holds. It does not tell you what the result means, which of its assumptions are load-bearing, or what should be built on it. That work is human, it is slow, and reporting at the time indicated the results had not yet been through peer review.
The same split is visible further down the field. In Quanta's August account of why the Erdos problems are falling to AI, Thomas Bloom, who maintains the canonical list of those problems, describes papers of one to two hundred pages that no human has read, produced by contributors not in a position to verify the output themselves.
That is verification debt.
One mechanism means this does not fix itself. In 2024, OpenAI researchers showed with prover-verifier games that when a model is optimized only to be correct, time-limited people get measurably worse at checking it, while training explicitly for legibility recovers much of that accuracy. Checkability is something you specify and pay for, and if it is not in the requirement it will not be in the product.
That is a procurement lever, and one of the few a buyer controls. Put verifier-legible output into your vendor evaluation criteria.
Article 14 of the EU AI Act already requires this
In the European Union human oversight is a legal obligation with specific content, and Article 50 transparency duties started applying earlier this month.
Article 14 of Regulation (EU) 2024/1689 requires that a high-risk system be provided so that the assigned overseers are enabled, as appropriate and proportionate, to do five things: understand its capacities and limitations and monitor its operation including detecting and addressing anomalies; remain aware of automation bias, the tendency to over-rely on the output; correctly interpret the output given the interpretation tools available; decide not to use it or to override it; and intervene or stop it.
Read those five against a tier-five output. A system whose reasoning nobody available can follow cannot be correctly interpreted, and anomalies inside it cannot reliably be detected. The regulation and the engineering reality arrive at the same conclusion from opposite directions.
One correction that changes who this lands on. Article 14 is a duty on the provider who builds the system. For most of what you buy, you are the deployer, and your duty is Article 26, which requires you to assign oversight to natural persons with the necessary competence, training and authority, and to ensure that oversight actually happens.
So the Act makes your vendor build the oversight capability into the product, and it makes you supply the competent humans to operate it. The second half is the half no vendor can discharge for you. That is the Verification Ladder's legal hook.
Whoever verifies, decides
If verification is the control point, then the people who verify set the defaults. Not in a philosophical sense. In the concrete sense of which outputs are allowed through, what counts as an acceptable answer, and which values are treated as hard constraints rather than preferences.
Two things are decided in two places. Local verification asks whether this output is correct, and building that is within your control. Upstream verification decides what the system will produce at all: what it refuses, what it treats as settled, what it presents without hedging. The first is bounded by the second, because you can only check what the system exposes to you.
Upstream is concentrated. The 2026 AI Index reports that industry produced more than 90 percent of notable frontier models in 2025, and that United States private AI investment reached $285.9 billion against $12.4 billion in China. Those figures measure concentration of model production and capital rather than of verification authority directly, and I am inferring the second from the first. The inference is that the teams who build the frontier models are also the teams who decide what those models treat as an acceptable answer, because today nobody else is positioned to.
This is a concentration claim, not a conspiracy claim. And the direction is not what a European reader assumes: the two officially claimed perfect scores in Shanghai came from Chinese companies, not from the American laboratories that dominate the investment figures. Concentration does not run in a single direction.
Here is the part I can defend and the part I cannot. Where no formal checker exists, and most commercial decisions have none, the only reliable error-catcher left is a second party who would have arrived by a different route. That is an argument about independence, and I think it holds. The step from there to "many legal, professional and cultural traditions should shape these defaults" is a further claim, and it is not the same claim. I believe different traditions surface different failure modes, so value plurality is also epistemic plurality. I cannot prove it. I do not think it is safe to compress that judgment into the review standards of a small number of laboratories, however capable those teams are.
The strongest case against me
The serious objection is not that I am reaching for villains. It is that concentrated verification might simply be better.
A small number of well-resourced expert institutions are more consistent, harder to shop between, and far cheaper to audit than a plural landscape. Plural standards are how regulatory arbitrage works: an actor shops for the loosest checker and calls it compliance. My own essay names that mechanism when it describes review being quietly loosened rather than resourced. And the counterweights I am about to endorse are convergence instruments, not plural ones.
The answer I have is narrow. Consistency is only a virtue once the standard is right, and a single tradition has no way to discover that its own standard is wrong, because the errors its starting assumptions make invisible have no external party who would notice. What I want is not many standards. It is one standard tested by people who did not write it.
I will also concede a number that does not help me. Documented AI incidents rose to 362 from 233 the year before, in the same year capability spread to more actors in more jurisdictions. That says the count tracks deployment rather than governance quality, which is exactly why an incident count is a poor proxy for whether anyone is checking.
The counterweights exist and are underused. The UNESCO Recommendation on the Ethics of Artificial Intelligence was adopted by 193 member states in November 2021 and states that AI systems should not displace ultimate human responsibility and accountability. The European Union has made oversight enforceable for high-risk uses. Neither specifies what counts as an acceptable answer. That question is still open, and it is being answered by default.
The counter-evidence
I am on the optimistic side of this. Cure the diseases. End the famine. Fix the climate accounting. Take poverty seriously.
So here is the counter-evidence against my own optimism.
In 2020 we had a shared problem, a legible problem, an urgent problem, and a working solution faster than almost anyone predicted. And we hoarded. Gavi's own assessment of COVAX attributes avoidable deaths to hoarding by wealthier countries, export controls, and manufacturing supply diversions and delays. COVAX delivered nearly 1.9 billion doses to 146 economies by the end of 2022. Its initial end-2021 target was some 950 million doses to lower-income countries; it delivered 830 million by that date, about 88 percent, and reached the target in mid-January 2022.
The capability was there. The distribution was a choice.
So what makes me think this time is different? Nothing does. Just hope.
There is one structural difference worth naming. In 2020 the bottleneck was manufacturing capacity and political will, and neither was inside my control or yours. This time the bottleneck is verification, and verification is something an organization can actually build. Hope is not a plan. Capacity is.
Salud, dinero y amor. And the fourth one.
Compliance is not the reason I care.
Salud, dinero y amor. Health, money and love. It is what you wish someone in Spanish when you raise a glass, and I have heard it my whole life.
I would add a fourth, because it is the one that actually decides things. Influence.
Health and wealth are, underneath everything, engineering and allocation problems: finding the molecule, and getting it to the person who needs it. Those are the two I think this technology genuinely reaches, and reaching either at scale would change more lives than any industrial program on record.
Love it cannot supply. It can make love more or less possible, which is not nothing, but it cannot supply it.
Influence it cannot create either. It can only move it.
So if this works, the argument will not be about the two it solved. It will be about the two it did not.
Influence is not an abstraction in this essay. It is the verification desk. Whoever sits there moves it.
The one thing to take from this
You will not out-know the model. Losing that contest is not a failure.
The job is to be the person who can tell whether it is right, and to make sure you are not the only kind of person in the room who is allowed to decide what right means.
Source and disclosure
Reporting on the 2026 International Mathematical Olympiad results and on the Erdos problems is linked in the text and is journalism rather than primary disclosure; the olympiad scores are company announcements that I have not been able to confirm against a jury statement. Everything else is drawn from the primary materials listed below. The verification ceiling, verification debt, and the Verification Ladder are my own framing and operating recommendation, not measured findings.
This article was developed with AI assistance and reviewed by Juan Beltrán. I selected the angle, verified the sources, challenged the claims, edited the argument, and accept editorial responsibility for the published version. The visuals were produced under human art direction. Figures in the evidence timeline were checked against the original sources on 24 August 2026. The two argument maps are conceptual and are labeled as such.
Written in a personal capacity. Views are my own. This is not legal, medical, or investment advice.