OpenAI’s sly mathematical breakthrough sends a chill through academia
OpenAI’s announcement Tuesday that it has solved one of mathematics’ legendary Millennium Prize problems should have been a moment of triumph. The result is both an undeniable achievement and a striking demonstration of just how rapidly AI is transforming mathematics. But before it was even formally announced, the breakthrough had been complicated by the unusual circumstances that prompted OpenAI to pursue the problem: After hearing other researchers were making progress, it seems to have thrown its considerable resources into a last-minute effort to beat them to the punch. The ensuing controversy has surfaced allegations of scooping, spying, and flagrant violations of long-standing academic norms that researchers fear could have a chilling effect on the field.
As Abhishek Saha, a mathematics professor at Queen Mary University of London, explains it, OpenAI has engaged in the “kind of things that mathematicians will generally not do.”
In a blog post published Tuesday, OpenAI said it took one of its unreleased models just 88 hours to find a solution to the Navier-Stokes problem, a thorny quandary concerning the movement of fluids. On account of the $1 million bounty available for whoever solves it, the problem is among mathematics’ most heavily researched, but it has nevertheless stumped human researchers for close to 90 years. OpenAI said its model solved the problem by focusing a swarm of roughly 10,000 AI agents powered by its internal model on the task and hailed the achievement as a “milestone.”
“If you don’t want me to be nice, then I don’t have to be nice.”
But the timing of the announcement has raised eyebrows. Just one day earlier, New York University mathematics professor Tristan Buckmaster published findings on a related problem with Levent Alpöge, a researcher at OpenAI’s archrival Anthropic (although Alpöge was not, here, working on behalf of his employer). Buckmaster said he contacted OpenAI after learning the company had become aware of their progress, to ask when it began working on the problem and what data its model had been trained on. The conversation, he said, quickly turned sour, with an OpenAI researcher asking him, “Why would you ruin your career?” when he said he would go public with what happened. When he asked why going public would ruin his career, Buckmaster said he received the following reply: “If you don’t want me to be nice, then I don’t have to be nice.” OpenAI urged Buckmaster to instead publish the work and credit OpenAI’s internal model, dropping Alpöge as coauthor.
Buckmaster said he asked OpenAI whether it had accessed his sessions on Codex, which he had used while tackling the problem, but that OpenAI grew increasingly evasive, even hostile, in its responses. In statements since, including the blog post announcing the result, OpenAI has flatly denied using any specific user data. “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem,” the company said.
But OpenAI could not conclusively rule out an indirect influence. “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models,” it said, while stressing that the two proofs differ significantly. Comments from OpenAI researchers on X echo those denials.
It is difficult to say exactly what happened. Timelines are tangled, research overlaps, and the provenance of AI-generated work is tough, if not impossible, to identify at the best of times. And it’s hardly a surprise that different players might compete to solve one of the most famous mathematical problems in the world, particularly one attached to a hefty prize.
But aspects of OpenAI’s account are hard to explain. By the company’s own telling, the effort was a hurried and incredibly expensive affair, costing it millions of dollars. Yet the company says it has no intention of claiming the bounty, which, in any case, has yet to be awarded by the Clay Mathematics Institute, which administers it. It said its only goal “is to report on the substantial progress of our AI models.” The company does not appear to have expended much effort on tackling Navier-Stokes before September, or if it has, it hasn’t spoken about it publicly.
OpenAI’s explanation effectively amounts to a thunderous “Why not?” The company said it began working on the problem after hearing rumors that other researchers were making progress on Millennium Prize problems. It found those rumors “on Twitter,” said OpenAI researcher Sébastien Bubeck at a press briefing reported on by Science. “So we thought to ourselves: ‘We have such a strong model. Why don’t we try to solve also a Millennium Prize problem?’” Bubeck said.
OpenAI said it only later realized the rumors concerned Alpöge and Buckmaster. Beyond addressing Buckmaster’s allegations about the use of his data, OpenAI has not publicly responded to his other claims and directed The Verge to its blog when asked for comment. Bubeck, who Buckmaster named in his account, has disputed parts of it, denying he ever asked Buckmaster to remove Alpöge as coauthor.
Even setting aside the most explosive allegations, aspects of OpenAI’s conduct the company has plainly acknowledged have shocked mathematicians. The apparent rush to beat other researchers to a result is simply not how mathematics is done in most cases. Scooping does happen, but it’s not easy, said Saha.
That’s partly because cutting-edge research often requires such deep and specialized expertise that few people are in a position to swoop in even if they wanted to, he explained.
Openness is a deeply embedded virtue in the discipline. “Mathematics depends heavily on an informal norm of trust,” said Matthew Ballard, a professor of mathematics at the University of South Carolina and associate director for scientific activities at the Institute for Computer-Aided Reasoning in Mathematics (ICARM). “Researchers routinely share incomplete ideas and ongoing work with colleagues to sharpen their thoughts. It is done with the expectation that it will not turn into a competition,” he said.
Though unable to comment on whether conversation logs might have been accessed, Carnegie Mellon professor Jeremy Avigad, who is also the director of ICARM, said that even “the thought that AI systems might steal ideas from our queries is chilling.” Mathematicians are accustomed to talking about their work without worrying about being scooped. “Now that even the slightest hint might be enough for someone with sufficient computational resources to set a swarm of agents on solving the problem, people are likely to be more cautious. It’s sad to think about how that might change the research environment.”
“Mathematics depends heavily on an informal norm of trust.”
OpenAI’s unwillingness or inability to say whether its models were informed by the work of other mathematicians compounds the sense of unease in the field. “That is a problem,” Brown University professor Brendan Hassett told The Verge, adding that “given the history of the AI companies appropriating copyrighted work without permission or payment, it is natural for people to ask these questions.” He said companies “should be held accountable to deliver” assurances that chat logs will not be used to improve their models. That includes being able to demonstrate that.
It’s unclear where exactly things go from here. Writing from a conference in Beijing, Yang-Hui He, a fellow at the London Institute for Mathematical Sciences, said he worries that “maths under the big companies is much too secretive.” As someone who says he is “always optimistic about AI,” he admits he is worried mathematics could be reverting to a more secretive state like in the past, when it was funded by patronage from wealthy families like the Medicis.
For most researchers, things may not change that much. There are only so many high-caliber problems companies like OpenAI and its rivals would be willing to spend such vast sums solving, Saha speculated. “You would not expect the AI labs to throw everything at most problems people work on because they just won’t get enough publicity.”
Publicity may have been part of the point, which could help explain why OpenAI decided to race after a problem it knew was connected to a researcher at Anthropic, even if he was acting independently. “This is clearly a PR victory for OpenAI,” said Oxford professor Andras Juhasz.
But Juhasz questioned how sustainable that approach could be, wondering whether this might spell the beginning of the end for AI companies’ involvement in research mathematics now that models can tackle some of the biggest problems. Human mathematicians scoop one another, too, he said, though what OpenAI did has shown that this can happen on a much grander scale. “Suddenly, 10,000 mathematicians jump on your problem,” he said.
All that could make OpenAI’s PR victory a Pyrrhic one. The company has, once again, proven that its models can compete at the very frontier of mathematics. In doing so, it appears to have alienated the very community it has been trying to impress.


