The math check-up

OpenAI just dropped over 700 'solutions' to some of the world's most brutal math problems, aiming to prove their models have big brain energy. They even consulted an advisory group—the Advisory Group on Mathematics and Artificial Intelligence (AGMAI)—to keep things legit. But real talk? The math community isn't buying the hype, and it’s giving major 'homework turned in without showing the work' vibes.

Why the experts are salty

AGMAI set some ground rules for how tech labs should handle these proofs, specifically asking them to stop using proprietary models and to ensure humans can actually understand the results. OpenAI basically ignored the first request and missed the mark on the second. Only a tiny fraction of their releases included the 'chain of thought' process, and less than half of the proofs were formalized—which is the industry standard for actually verifying that the math is correct.

Prominent mathematician Terence Tao dragged the current approach on social media, pointing out that AI prompters are just hunting for a 'solved' stamp without actually understanding the results or engaging with the field. It’s a huge L for academic integrity when you can’t answer questions about the proof you just 'wrote.'

The translation fail

A new paper from researchers at Cambridge and King’s College London highlighted a major red flag: the gap between the AI’s 'natural language' explanations and the formal code it uses to verify them. They found discrepancies in a solution OpenAI claimed for a massive Navier-Stokes equation problem. Essentially, the model is 'lost in translation,' and experts say we shouldn't be trusting these outputs without serious peer review. It’s wild that they're skipping the human sanity check that’s supposed to be the backbone of scientific discovery.

Why it matters

Math isn't just about getting a result; it's about the knowledge gained that helps solve other problems. When AI spits out a solution that no human understands, the work hasn't been finished—it's actually just beginning. As Harvard professor Melanie Wood put it, there’s no human understanding at the point of release, making these big claims feel more like marketing than a scientific breakthrough.