AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI Can Do The Work Cheaply. Review Still Takes Resources on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source essay argues that AI is reducing the cost of producing mathematical manuscripts, code and contract work while leaving human review constrained. The figures it cites point to a growing verification bottleneck, though several software metrics come from vendors and should be read with that limitation in mind.

An analysis published this week argues that AI is making work faster and cheaper to produce without making it equally easy to verify, using examples from mathematical research, software development and contract work. Its central concern is that human review capacity may become a constraint as AI-generated output grows, although the evidence varies in strength and some software figures come from companies that sell review tools.

The analysis says OpenAI published 722 mathematical manuscripts this week, with an average result taking about three hours of compute. It says the model was given about 4,000 problems and the manuscripts covered 372 families. Some results were formally checked in Lean, while OpenAI warned that unformalized work could contain issues. The source contrasts this production scale with the careful verification of an earlier counterexample to an Erdős conjecture by five prominent mathematicians.

In software, the analysis cites Faros AI data showing teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. It also cites a LinearB analysis of 8.1 million pull requests across 4,800 organizations: AI-generated changes reportedly waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study cited in the piece found that 61% of AI-agent pull requests received no human review before being merged or closed. These measures cover different datasets and should not be treated as interchangeable.

A third example is a partnership between OpenAI and contract-software company Ironclad. The source says GPT-6 Astra was evaluated on 11 contracting tasks and met 55% of the evaluation criteria on average, an improvement over its predecessor. That result indicates progress, but the remaining criteria still matter in professional use: missed approval requirements or an incorrect jurisdiction clause may need to be caught before a draft is relied on.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA source essay links AI-driven growth in mathematical, software and contract output to persistent limits on human verification.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Shapes AI’s Value

If AI produces more drafts, code and research than specialists can check, organizations may not be able to use that output safely or efficiently. The analysis describes three possible responses already reflected in its examples: less review, as with pull requests that receive no human check; slower review, as AI-generated changes wait in queues; and greater reliance on the producer’s own selection of what to publish.

Those responses carry different risks. Skipped checks can let defects through, while reviewers who assume machine-generated work is unreliable may delay useful contributions along with flawed ones. And when the producer chooses which results merit attention, independent scrutiny can become thinner. The piece’s figures do not establish how widespread each outcome is across all organizations, but they show why output volume alone is a poor measure of usable productivity.

The analysis also raises a workforce concern: reviewers need experience. If junior staff do less first-draft work because AI supplies it, they may get fewer opportunities to develop the judgment required for senior review. The author argues that organizations should preserve training pathways alongside AI adoption. That is an interpretation of the likely long-term effect, not a demonstrated outcome in the cited data.

Amazon

AI verification tools for software review

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Bottleneck

The examples span tasks where correctness has different meanings. A proof assistant can check whether a formal proof follows from a stated theorem, but a human still may need to judge whether the theorem addresses a meaningful question. Software tests check specified behavior; they cannot establish that the tests cover every user need or failure. Contract review also involves rules and responsibilities that can depend on jurisdiction and business circumstances.

The source frames this gap as verification abundance and adjudication scarcity: tools may test a defined answer, while people decide whether the question, assumptions and consequences are appropriate. It also points to accountability. A person or institution may be responsible for a signed contract, an engineering design or a research paper; the analysis argues that this responsibility cannot simply be assigned to a model.

The software statistics deserve particular care. Faros AI and LinearB sell products related to engineering workflow or code review, according to the source, which says their commercial interests are a reason to scrutinize the figures. The piece also cites peer-reviewed research, but its summary does not provide the study’s methods or publication details. The measurements therefore offer evidence of pressure on review, not a single definitive estimate of its scale.

Amazon

mathematical manuscript verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Missing?

The available figures do not establish whether the reported patterns apply across industries, company sizes or types of AI system. The source gives different observation periods and methods for the software data, and does not provide enough detail here to compare them directly. It is also unclear how often unreviewed work causes consequential failures, or whether teams later inspect changes that initially bypassed review.

The contracting evaluation’s 55% figure is an average across 11 tasks, but the source does not specify the evaluation criteria, how the score was calculated or how performance varied by task. Nor does it establish that all remaining criteria represent errors of equal severity. The analysis’s concern about reduced apprenticeship is a plausible risk it raises, but no measured trend in training or career outcomes is supplied.

Amazon

contract review AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking the Human Review Queue

The next useful evidence will show whether organizations can scale review alongside generation: for example, by reporting review times, rates of human inspection and error outcomes over comparable periods. Details about the methods behind the cited studies, and whether their findings hold across different teams, would help establish how broad the bottleneck is.

For now, the analysis points to a practical question for employers: whether AI adoption is paired with enough qualified reviewers and with training that helps newer workers build judgment. The source does not announce a new policy or a scheduled follow-up; the longer-term effects on productivity, quality and professional development remain to be measured.

Amazon

AI code review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The source argues that AI is expanding the production of research, code and contract work faster than people can verify it. It presents this as a developing review-capacity problem, not as proof that all AI output is unreliable.

How many mathematical manuscripts does the source say OpenAI published?

It says OpenAI published 722 manuscripts this week, drawn from work on about 4,000 problems. The source says some results were formally checked in Lean and that unformalized results could have issues.

What do the software figures show?

The cited analyses report more pull requests alongside longer or less frequent review. For example, LinearB’s reported figures say AI-generated changes waited 4.6 times longer for review to begin and had a lower acceptance rate than human-written changes. The data comes from specific datasets and does not by itself describe every software team.

Can AI systems review other AI-generated work?

They can help with defined checks, such as testing code against specified requirements. The analysis argues that people may still need to judge whether those requirements are correct, whether important risks were missed and who is accountable for the final work.

What remains unknown?

The reach of the reported patterns, the consequences of skipped review and the long-term effect on professional training are not established by the figures summarized in the source. More comparable data on review quality and outcomes is needed.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Monster Hunter Wilds Climbing The Steam Charts

Monster Hunter Wilds has climbed to the top of Steam’s most-played games, reaching a peak of over 42,600 players. The trend signals rising interest in the game.

xAI Launches Imagine Image 2.0 In Grok Quality Mode – TestingCatalog AI News

xAI has introduced Imagine Image 2.0 within Grok’s Quality Mode, but technical details, availability, and performance remain unconfirmed.

Can Anthropic’s Claude Startups Program Win Over Fast-Growing Companies?

A CNBC report headline says Anthropic is expanding Claude Startups to attract founders and fast-growing companies, but gives no terms or timeline.

Commodore 64 Released September 1, 1982

The Commodore 64 was officially launched on September 1, 1982, marking a significant milestone in home computer history and consumer electronics.