OpenAI’s Ten Math Proofs All Pass the Machine Check. Two Mathematicians Say the Ideas Were Already Theirs.


a blackboard with a lot of writing on it

Every one of the ten proofs OpenAI published on August 1 compiles in Lean 4. Feed them to the checker and it returns green: the logic holds, line by line, with no human required to trust it. That was supposed to be the end of the argument. Nine days later, two of the mathematicians whose prior work sits inside those proofs say the checker verified the wrong thing.

On August 6, Scientific American reported that named researchers are accusing OpenAI of research misconduct over the ten-result math report the company used to introduce its next model family, Astra. The dispute is not about whether the proofs are correct. It is about whether they are original, and whether the people who did the earlier work got credit. Those are two different questions, and the machine can only answer one of them.

What a Lean certificate actually guarantees

When OpenAI published the Astra results, the formalization in Lean was the headline improvement over the last time. In May, when OpenAI claimed a result on the Erdős unit distance problem, the scrutiny landed on whether the proof held up at all. This site wrote at the time that the Lean certificates “close that gap.” They do. A proof that compiles in Lean 4’s kernel has been mechanically verified as logically sound. Anyone can rerun the check and get the same binary verdict.

The mistake, mine included, was treating that verdict as the whole trust problem. Lean answers one narrow question with total certainty: given these premises, does this chain of inference reach this conclusion without a gap? It says nothing about where the premises came from, whether the central idea is new, or whether an identical argument was published by someone else years earlier. A plagiarized proof compiles exactly as cleanly as an original one. The checker has no notion of authorship, and it was never designed to.

That distinction is the entire story here, and it is the part almost every write-up glossed over on August 2, including this one.

The two disputed results

Steven Miller, a mathematician at Yeshiva University, says one of the sphere-packing results presents an argument as OpenAI’s own when it first appeared in a 2016 paper he co-authored. He is not calling it a citation slip. “They are running roughshod over the work of others who came before them in a deliberate way,” he told Scientific American. “It seems completely systematic to me, and it points to research misconduct.” That is the strongest charge in the piece, and Miller frames it as intent, not oversight.

The second dispute is more revealing precisely because it is softer. The soficity result, which OpenAI presented as constructing the first non-sofic groups, is one that group theorists had circled for years. Francesco Fournier-Facio, a group theorist at the University of Cambridge, says the “breakthrough” stitched together ideas already present in papers from 2016 and 2019, and he was able to reconstruct the argument himself. His objection is aimed at the framing, not the math: “there is the big PR machine that wants to sound as impressive as possible and does not care about being 100 percent accurate.” Andreas Thom of Dresden University of Technology, a co-author of the 2019 work being drawn on, described the result as “creative and at the same time elementary.” Even the aggrieved party granted it some creativity.

Hold those two reactions side by side, because the gap between them is where the useful signal lives. One mathematician alleges deliberate plagiarism. Another says the underlying work is real and even elegant, but the packaging oversold its novelty. Both can be true at once, and neither is the clean “the AI made it up” story that travels well. The Lean check is irrelevant to both complaints. The proofs are correct. That was never in question this round.

OpenAI’s response, and an inverted defense

OpenAI’s on-record answer was procedural. “We take responsibility for the correctness of these results and are meeting the same standards generally expected of human mathematicians,” a spokesperson said, adding that the company planned “small updates” to the paper that week, “consistent with standard academic practice.” The company also quietly walked back its original claim that the problems had “seen no progress on the main result for at least a decade,” softening the language after the objections surfaced. Correctness, again, is the ground OpenAI wants to stand on, and it is solid ground. It is also not the ground under dispute.

The stranger part is how OpenAI reportedly invoked the Leiden Declaration in its defense. That declaration, published June 2 and endorsed by the International Mathematical Union with more than 3,000 signatories including Terence Tao and Peter Scholze, was written as a warning to AI companies. It lists five risks: unreliable results, missing citations, dependence on closed commercial systems, exaggerated claims, and loss of scientific independence. It explicitly called out labs that announce results through press releases instead of peer-reviewed journals. Per The Decoder, OpenAI cited that same document to argue the credit belongs to the AI rather than to a human author, reframing an attribution complaint about uncredited prior researchers into a philosophical point about whether the model deserves a byline. A declaration written to protect human mathematicians from exactly this behavior got repurposed as a shield for it.

This is a pattern, not a stumble

The reason the OpenAI math plagiarism dispute matters is that this is the second time in under a year the company has announced a math result through a press release, skipped peer review, and then had specialists dispute the framing. In May, mathematicians pointed out that an earlier “solved in 24 hours” conjecture had every sentence technically correct while the result drifted away from the original claim. The Lean formalization was OpenAI’s answer to that criticism. It closed the correctness hole and left the provenance hole wide open, because those were always separate holes.

There is a structural incentive underneath the pattern. A frontier lab racing toward what OpenAI now describes as a research-grade autonomous system has every reason to present its model’s output as maximally novel and maximally unassisted by prior human work. Novelty is the product. Careful attribution of the 2016 and 2019 papers the argument leans on makes the achievement look smaller, and “our model rediscovered a known technique and applied it well” is a far less fundable headline than “our model solved ten open problems.” The verification theater, the compiling Lean proofs, provides cover, because it lets the company answer a correctness challenge it will win while the harder attribution challenge goes unaddressed.

Why a green checkmark is not a clean bill

Step out of pure mathematics and the lesson generalizes to anyone putting AI output into work they will sign. The industry has spent 2026 building verification layers precisely because model output cannot be taken on faith, and those layers are worth having. But it is worth being clear-eyed about what each one certifies. A test suite confirms code runs; it does not confirm the code is not copied from a license-incompatible repository. A Lean proof confirms a theorem holds; it does not confirm the theorem is yours. The spreadsheet audit problem this site covered on August 10 is the same shape from the other direction: a financial model whose totals foot perfectly while a subtly wrong line item hides inside them. Correct-looking is not the same as correct, and correct is not the same as original or properly sourced.

For enterprise buyers, that maps to a governance gap most procurement checklists do not yet have a box for. When a model drafts a research memo, a market analysis, or a body of code, the natural instinct is to ask whether it is accurate. The instinct that is missing is asking whether it is derivative in a way that creates legal or reputational exposure, and whether the sources it built on are acknowledged. This site has argued repeatedly that a benchmark number tells you almost nothing until you build your own eval on your own data. The attribution problem is the version of that principle you cannot benchmark at all: no automated grader flags an uncredited source, because recognizing prior art requires knowing the prior art existed.

The Astra math results are, by the machine’s measure, correct. That is a real accomplishment and it should not be waved away by the controversy. But the past nine days have drawn a line that anyone deploying AI in serious work should internalize. Verification and provenance are different guarantees, and the tools that give you the first one confidently give you nothing on the second. OpenAI built a system that can produce a proof no human needs to trust. It has not built one that can tell you whether that proof was ever really its own.

Ty Sutherland

Ty Sutherland is the Chief Editor of AI Rising Trends. Living in what he believes to be the most transformative era in history, Ty is deeply captivated by the boundless potential of emerging technologies like the metaverse and artificial intelligence. He envisions a future where these innovations seamlessly enhance every facet of human existence. With a fervent desire to champion the adoption of AI for humanity's collective betterment, Ty emphasizes the urgency of integrating AI into our professional and personal spheres, cautioning against the risk of obsolescence for those who lag behind. "Airising Trends" stands as a testament to his mission, dedicated to spotlighting the latest in AI advancements and offering guidance on harnessing these tools to elevate one's life.

Recent Posts