Ten problems that had stumped mathematicians for years, some of them for decades, fell in a single batch last week. The machine that solved them has not been released, does not have a price, and does not yet have a confirmed name beyond the one OpenAI gave it: Astra.
On Friday, OpenAI said an internal version of Astra, which it describes as its next major model family, had solved or made substantial progress on ten open questions across mathematics, quantum complexity, and theoretical computer science. Noam Brown, a research scientist at the company, announced the results and called them a major step forward for scientific reasoning. Sam Altman had spent part of the week in Washington demoing the model to political staffers, which meant a room in DC saw what Astra could do before the rest of us got a look.
What Astra actually solved
The list is not light reading. The problems touch high-dimensional sphere packing, binary and spherical codes, arithmetic circuit complexity, quantum parallel repetition, the closest vector problem, Ehrhart’s volume conjecture, and multicolour Ramsey numbers. The result drawing the most attention is an explicit construction of a non-sofic group, a question that had been open since Mikhail Gromov introduced the idea of soficity in 1999. For a quarter of a century nobody could show such a group existed. OpenAI says its model wrote one down.
Two details separate this from the usual round of benchmark bragging. The first is cost. OpenAI disclosed that the tokens needed to produce all ten solutions would run to roughly $2,000 at its Sol API rates, an average of about $200 a problem. For work that would keep specialist researchers busy for months, that number is small enough to make people uneasy.
The second is verification. Every result shipped with a Lean 4 certificate. Lean is a proof assistant, a programming language a mathematician can use to encode a logical argument so a computer can check each step. OpenAI posted the certificates to GitHub along with a manuscript collection running to hundreds of pages and walkthroughs of the model’s reasoning. The certificates reportedly compiled with a “sorry” count of zero, meaning no step was left unproven and quietly skipped. A machine either passes Lean’s kernel or it does not, and no doctorate is required to read the verdict.
Why the $2,000 comes with an asterisk
The price tag is real, but it counts only the wins. It leaves out the compute OpenAI spent on problems the model failed to crack, and Brown was candid that there were failures. He noted the team did not spend much on each individual problem and that test-time compute could be pushed a great deal further. Read generously, that is a claim the ceiling sits well above what we just saw. Read skeptically, it is a reminder that a curated list of successes is exactly what a company tends to publish about a product it has not shipped.
There is precedent for the caution. Back in May, OpenAI said an unreleased model had made progress on the Erdos unit distance conjecture, a problem roughly eighty years old. That claim arrived without a cost figure and without formal proofs a mathematician could run. This time the company brought receipts. That is a real step up, even for readers inclined to discount the announcement.
The harder problem is judging the news
Here is the strange part. Most people, including most journalists, have no way to independently assess whether constructing a non-sofic group is a minor curiosity or a landmark. The mathematics sits far outside ordinary expertise, and even working mathematicians tend to specialise narrowly enough that one field’s breakthrough reads like another’s foreign language.
So a pattern has emerged that says as much about the moment as the math does. When researchers wanted to gauge how hard these ten problems really were, some of them asked a different AI to rate the difficulty. The answers came back suggesting that any single one could plausibly anchor a serious case for a Fields Medal, mathematics’ highest honour. Whether or not that holds up, the method is the tell. We are starting to evaluate what AI can do by asking AI, because the humans qualified to check the work are getting scarce relative to the pace of the claims.
OpenAI has not given Astra a release date, or said whether it will arrive as a point update or a new flagship. What it has done is shift the argument. The question of whether these systems can produce real, original, checkable mathematics is quieter than it was a week ago. The question of what that means, and who is left qualified to referee it, is only starting. For more coverage of frontier AI models, visit Mylistingo.
Source: How Significant Are AI’s Latest Math Breakthroughs?, The AI Daily Brief.







