🔍 Read the full analysis: AI Can Do The Work Cheaply. Review Still Takes Resources on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source essay argues that AI is reducing the cost of producing mathematical manuscripts, code and contract work while leaving human review constrained. The figures it cites point to a growing verification bottleneck, though several software metrics come from vendors and should be read with that limitation in mind.
An analysis published this week argues that AI is making work faster and cheaper to produce without making it equally easy to verify, using examples from mathematical research, software development and contract work. Its central concern is that human review capacity may become a constraint as AI-generated output grows, although the evidence varies in strength and some software figures come from companies that sell review tools.
The analysis says OpenAI published 722 mathematical manuscripts this week, with an average result taking about three hours of compute. It says the model was given about 4,000 problems and the manuscripts covered 372 families. Some results were formally checked in Lean, while OpenAI warned that unformalized work could contain issues. The source contrasts this production scale with the careful verification of an earlier counterexample to an Erdős conjecture by five prominent mathematicians.
In software, the analysis cites Faros AI data showing teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. It also cites a LinearB analysis of 8.1 million pull requests across 4,800 organizations: AI-generated changes reportedly waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study cited in the piece found that 61% of AI-agent pull requests received no human review before being merged or closed. These measures cover different datasets and should not be treated as interchangeable.
A third example is a partnership between OpenAI and contract-software company Ironclad. The source says GPT-6 Astra was evaluated on 11 contracting tasks and met 55% of the evaluation criteria on average, an improvement over its predecessor. That result indicates progress, but the remaining criteria still matter in professional use: missed approval requirements or an incorrect jurisdiction clause may need to be caught before a draft is relied on.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Shapes AI’s Value
If AI produces more drafts, code and research than specialists can check, organizations may not be able to use that output safely or efficiently. The analysis describes three possible responses already reflected in its examples: less review, as with pull requests that receive no human check; slower review, as AI-generated changes wait in queues; and greater reliance on the producer’s own selection of what to publish.
Those responses carry different risks. Skipped checks can let defects through, while reviewers who assume machine-generated work is unreliable may delay useful contributions along with flawed ones. And when the producer chooses which results merit attention, independent scrutiny can become thinner. The piece’s figures do not establish how widespread each outcome is across all organizations, but they show why output volume alone is a poor measure of usable productivity.
The analysis also raises a workforce concern: reviewers need experience. If junior staff do less first-draft work because AI supplies it, they may get fewer opportunities to develop the judgment required for senior review. The author argues that organizations should preserve training pathways alongside AI adoption. That is an interpretation of the likely long-term effect, not a demonstrated outcome in the cited data.
AI verification tools for software review
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Bottleneck
The examples span tasks where correctness has different meanings. A proof assistant can check whether a formal proof follows from a stated theorem, but a human still may need to judge whether the theorem addresses a meaningful question. Software tests check specified behavior; they cannot establish that the tests cover every user need or failure. Contract review also involves rules and responsibilities that can depend on jurisdiction and business circumstances.
The source frames this gap as verification abundance and adjudication scarcity: tools may test a defined answer, while people decide whether the question, assumptions and consequences are appropriate. It also points to accountability. A person or institution may be responsible for a signed contract, an engineering design or a research paper; the analysis argues that this responsibility cannot simply be assigned to a model.
The software statistics deserve particular care. Faros AI and LinearB sell products related to engineering workflow or code review, according to the source, which says their commercial interests are a reason to scrutinize the figures. The piece also cites peer-reviewed research, but its summary does not provide the study’s methods or publication details. The measurements therefore offer evidence of pressure on review, not a single definitive estimate of its scale.
mathematical manuscript verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Missing?
The available figures do not establish whether the reported patterns apply across industries, company sizes or types of AI system. The source gives different observation periods and methods for the software data, and does not provide enough detail here to compare them directly. It is also unclear how often unreviewed work causes consequential failures, or whether teams later inspect changes that initially bypassed review.
The contracting evaluation’s 55% figure is an average across 11 tasks, but the source does not specify the evaluation criteria, how the score was calculated or how performance varied by task. Nor does it establish that all remaining criteria represent errors of equal severity. The analysis’s concern about reduced apprenticeship is a plausible risk it raises, but no measured trend in training or career outcomes is supplied.
As an affiliate, we earn on qualifying purchases.
Tracking the Human Review Queue
The next useful evidence will show whether organizations can scale review alongside generation: for example, by reporting review times, rates of human inspection and error outcomes over comparable periods. Details about the methods behind the cited studies, and whether their findings hold across different teams, would help establish how broad the bottleneck is.
For now, the analysis points to a practical question for employers: whether AI adoption is paired with enough qualified reviewers and with training that helps newer workers build judgment. The source does not announce a new policy or a scheduled follow-up; the longer-term effects on productivity, quality and professional development remain to be measured.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The source argues that AI is expanding the production of research, code and contract work faster than people can verify it. It presents this as a developing review-capacity problem, not as proof that all AI output is unreliable.
How many mathematical manuscripts does the source say OpenAI published?
It says OpenAI published 722 manuscripts this week, drawn from work on about 4,000 problems. The source says some results were formally checked in Lean and that unformalized results could have issues.
What do the software figures show?
The cited analyses report more pull requests alongside longer or less frequent review. For example, LinearB’s reported figures say AI-generated changes waited 4.6 times longer for review to begin and had a lower acceptance rate than human-written changes. The data comes from specific datasets and does not by itself describe every software team.
Can AI systems review other AI-generated work?
They can help with defined checks, such as testing code against specified requirements. The analysis argues that people may still need to judge whether those requirements are correct, whether important risks were missed and who is accountable for the final work.
What remains unknown?
The reach of the reported patterns, the consequences of skipped review and the long-term effect on professional training are not established by the figures summarized in the source. More comparable data on review quality and outcomes is needed.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
