OpenAI says its newest internal AI model has resolved more than 100 long-standing open problems across mathematics. If even independent mathematicians have confirmed many of those results, researchers could face a new problem: AI may produce research faster than people can check it, understand it, and decide what it contributes.
The model began training on August 28, based on OpenAI’s September 21 announcement. OpenAI says its progress surprised the company’s own mathematicians. The same model also produced OpenAI’s recently announced solution to the Navier-Stokes Millennium Prize problem.
OpenAI has not released most of the 100-plus results publicly, so outside researchers cannot yet judge their quality or novelty. It's now working with an independent group of mathematicians on how to review and release those results.
The unusual part is that OpenAI says it has produced more than 100 results, but it has not yet released most for outside review.
Finding a proof may become faster than understanding it
When an AI produces a proposed proof, mathematicians have more work to do than checking whether each logical step is valid. They also need to confirm that the proof answers the intended problem, determine whether someone has already established part of the result, and understand which ideas are actually new.
OpenAI’s Navier-Stokes result provides an early example. The company published both a written proof and a version in Lean, software that can mechanically check mathematical reasoning against defined assumptions.
OpenAI says the successful research effort used roughly 10,000 concurrent agents with access to cached internet material and code execution. The system generated the proposed solution in about 88 hours, followed by roughly 17 hours of Lean formalization.
Verification can speed up part of the review process because mathematicians do not have to check every logical step. But as The Neuron previously covered, human reviewers still need to confirm a verified proof addresses the original mathematical problem.
Researchers also need to understand why a proof works. A correct result becomes more valuable to the field when mathematicians can identify the ideas behind it and apply those ideas elsewhere.
If AI starts producing results faster than researchers can complete those steps, mathematical research could develop a review backlog. The bottleneck would move from finding possible answers toward deciding which answers deserve attention and what researchers can learn from them.
Mathematicians are already debating what counts as progress
Some mathematicians argue that counting solved problems gives an incomplete picture of research progress.
In the September 11 declaration “A Severe Misalignment of AI in Mathematics,” mathematicians, including Fields Medal winners, criticized using famous open problems as benchmarks for AI systems. They argue that mathematical progress also depends on understanding why a result works and whether its ideas can support further research.
A company can announce that its model solved an open problem relatively quickly. The research community, however, must examine the result, compare it with earlier work, resolve questions of credit, and explain the mathematics clearly enough for other researchers to use it.
OpenAI now has to decide how quickly to release the results
OpenAI has turned to the independent Advisory Group on Mathematics and Artificial Intelligence for help managing its reported backlog of results.
The group, hosted at the Institute for Advanced Study, says its current task is to advise OpenAI on how to coordinate the release of a large number of mathematical results produced by the company's internal model. Its members are unpaid and can publish their recommendations independently.
The group does not have decision-making authority over OpenAI, however, and its role does not include controlling how quickly the company continues its internal mathematical research.
OpenAI could therefore keep generating results even if publication proceeds more slowly. A controlled release schedule may give mathematicians more time to review each result, but it does not reduce the rate at which new work enters OpenAI's internal queue.
Faster AI research could create more work for human experts
The same problem could appear beyond mathematics. If AI systems produce research outputs faster, expert evaluation may consume a larger share of the total research process.
Mathematics provides an unusual test because proof systems can automate part of verification. Other scientific fields may require experiments, physical measurements, or other forms of evidence that take much longer to produce.
OpenAI's claim of over 100 resolved problems is still a claim until researchers can inspect the work. As those results come out, the number of correct proofs will be only part of what mathematicians learn.
The other measure will be how long it takes people to understand those proofs well enough to decide what they add to mathematics. If AI can generate discoveries faster than researchers can absorb them, producing new results may become the faster part of scientific research.