OpenAI announced this week that one of its AI models cracked a Millennium Prize problem in 88 hours. If you're keeping score at home, that's one of seven math problems carrying a million-dollar bounty, the kind of challenges that have stumped the world's best mathematicians for decades. The Clay Mathematics Institute set them up in 2000 as modern successors to Hilbert's problems. Only one has been solved so far, and that took Grigori Perelman seven years and a withdrawal from public life.
So when OpenAI says its model did it in under four days, you'd expect fanfare. Instead, it walked straight into a fight.
The problem isn't whether the AI produced something. It did. The question is whether what it produced counts as a solution in the way mathematicians actually use that word. OpenAI's release was light on proof structure and heavy on runtime metrics. Meanwhile, mathematicians on social media and in academic circles have been less than impressed, pointing out that claims of solving Millennium problems require peer review, formal verification, and a level of rigor that a statistical model grinding through token predictions doesn't automatically provide.
This isn't just about ego or gatekeeping. It's about what we're willing to call "solved." Mathematics has spent centuries building a system where a proof isn't valid until it's been checked, reproduced, and survives scrutiny from people who spend their lives in that problem space. You don't get to skip that process just because your model has a lot of parameters.
The timing is also telling. OpenAI has been pushing hard on the idea that AI can do real science, not just autocomplete or summarization. That's a useful narrative when you're raising billions and asking for regulatory patience. But science has referees. Math has standards. And if your big breakthrough can't survive contact with the people who actually work in the field, you didn't solve the problem—you generated something that looked like a solution to your eval harness.
There's a version of this story where AI becomes a genuine tool for mathematical discovery. But it starts with submitting to the process, not skipping it. Until OpenAI publishes the proof, submits it for review, and lets the mathematical community do what it does, this is a press release, not a milestone. And mathematicians, apparently, are not impressed by press releases.