🔍 Read the full analysis: AI Made It Easier To Do The Work. Harder To Know It’s Right on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A source article describes a widening gap between the speed of AI-generated work and the human capacity to verify it, citing mathematics, software development and contract work. The figures point to a potential review bottleneck, but some software data comes from vendors, and the long-term effects on jobs and training remain uncertain.
AI systems are producing mathematical manuscripts, software changes and contract drafts faster than human experts can verify them, according to an analysis by ThorstenMeyerAI.com. Its examples span three fields and suggest that review capacity—not generation—may be limiting how much AI-produced work organisations can safely use.
The analysis says OpenAI’s model was given about 4,000 mathematical problems and generated 722 manuscripts grouped into 372 families. It reports that the average result took about three hours of compute. Some manuscripts were formally checked in Lean, a proof-assistant system; OpenAI cautioned that some results without formal verification “could have issues.” The source contrasts that volume with the careful scrutiny given to an earlier result from the programme, described as a counterexample to an Erdős conjecture.
In software, the analysis cites several datasets that report more code changes but heavier or less consistent review. Faros AI found teams merged 98% more pull requests in high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study, as summarized by the source, found that 61% of AI-agent pull requests received no human review before being merged or closed.
The analysis also points to OpenAI’s partnership with contract-software company Ironclad. It says GPT-6 Astra, trained on real contracting workflows, met an average of 55% of evaluation criteria across 11 tasks. That is described as an improvement over an earlier model, but the source does not provide the earlier score or details of the evaluation. The figures suggest that AI can assist with professional work without establishing that its output is ready to use without expert checking.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity May Limit AI Use
If production becomes faster while review remains slow, organisations may not be able to use every result their AI tools generate. The constraint can affect delivery speed, quality and accountability: work may wait for a qualified reviewer, be accepted with less scrutiny, or be set aside because checking it costs too much.
The source describes three possible responses already visible in the cited examples: work is merged without review, reviewers deprioritise machine-generated changes, or producers decide which outputs deserve attention. Each has a different cost. Skipped checks can leave mistakes undiscovered; blanket suspicion can delay good work; and relying on the producer’s own selection means an independent reviewer may never assess what was left out.
The analysis also raises a workforce concern. In software, law and research, experienced reviewers often develop their judgement by doing the underlying work themselves. If entry-level workers mainly supervise machine-generated drafts rather than learning to produce them, organisations could weaken the pipeline of future experts. That is a concern, not an established outcome; the source offers no long-term workforce data to show whether this is happening at scale.
AI verification tools for software development
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Review Pressures
Mathematics provides a clear distinction between checking a formal object and judging its meaning. A proof assistant can verify that a proof follows from stated assumptions and establishes a stated theorem. It cannot decide on its own whether the theorem addresses the intended question, whether the result is important, or what it contributes to the field. The source summarizes this tension as “verification abundance, adjudication scarcity.”
Software has measurable review activity, but the cited statistics need qualification. The source notes that several organizations supplying the figures sell code-review tools, which gives readers reason to examine methods and definitions carefully. The metrics also describe different measures: review wait time, acceptance rates, review coverage and pull-request volume are not interchangeable. Even so, the source says the studies point in a similar direction: more code is being produced, while review remains a constraint.
The Ironclad example broadens the issue beyond technical teams. Contracts and other professional documents can contain requirements tied to jurisdiction, approvals or organizational policy. A model’s performance against an evaluation is evidence about those tasks, but it does not establish that a draft is legally suitable in every case. People and institutions still carry responsibility for decisions made using the work.
mathematical proof assistant software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Reliable Are the Comparisons?
The source does not supply links to the underlying datasets, study methods or full evaluation results, so the reported figures cannot be independently assessed from the material provided here. In particular, the software statistics come from different organizations and may use different definitions of AI-generated work, review time and acceptance. Several cited sources sell code-review products, a potential commercial interest the analysis itself flags.
It is also unclear how representative the examples are across organizations, job types or AI systems. The source does not establish whether the observed review delays or missing reviews were caused by AI adoption, nor whether teams later improved their processes. The reported 55% score for Astra lacks a stated baseline and does not show how the 11 contract tasks were weighted or how performance translates to real-world use.
Finally, the longer-term effects on hiring and professional training remain predictions. The material argues that reduced practice could weaken future reviewer supply, but provides no longitudinal evidence confirming that outcome. The scale and pace of any “referee premium” for experienced reviewers are also unspecified.
As an affiliate, we earn on qualifying purchases.
More Evidence on Review Quality
The next useful evidence would include independent, comparable studies that track not only how much AI-generated work is produced, but also how often it is checked, what errors reviewers find and what happens after work is accepted. For mathematics, that means distinguishing formal proof verification from expert assessment of a result’s relevance and assumptions. For software and contracts, it means reporting methods and outcomes in enough detail for teams to compare results across settings.
Organisations adopting these tools will also need to decide who is accountable for approving outputs and how junior staff can gain the experience needed to make those decisions. The source argues for protecting training pathways, but does not describe a settled model for doing so. Whether review tools, formal checks or new staffing practices can close the gap remains an open question.
contract review automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The analysis says AI is making it faster and cheaper to produce work in fields including mathematics and software, while human verification remains comparatively slow and limited.
Did OpenAI’s model formally prove all 722 mathematical results?
No. The source says some manuscripts were formally checked in Lean and quotes OpenAI warning that some unformalized results “could have issues.” It does not say that every result was formally verified.
What do the software figures show?
The cited sources report higher pull-request volumes alongside longer review waits, lower acceptance rates for AI-generated changes in one dataset and substantial shares of agent pull requests without human review. The measures come from different studies and should not be treated as one directly comparable dataset.
Does the analysis prove AI is reducing entry-level training?
No. It raises the possibility that junior workers could get less practice if AI performs more drafting and coding, but the supplied material does not include long-term evidence establishing that effect.
What remains uncertain about the contract-software example?
The source reports that GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, but gives no earlier-model score, detailed task results or evidence that the score predicts performance in live legal work.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
