Back to list
Aug 03, 2026
8
0
0
GPTNEW

OpenAI Astra Solves Ten Open Math Problems, Lean-Verified

OpenAI unveiled Astra, its next model family, with ten machine-checked solutions to open math and CS problems, all formalized in Lean.

#OpenAI#GPT#Astra#Lean#Mathematics
OpenAI Astra Solves Ten Open Math Problems, Lean-Verified
AI Summary

OpenAI unveiled Astra, its next model family, with ten machine-checked solutions to open math and CS problems, all formalized in Lean.

Introduction

On August 1, 2026, OpenAI published a research report titled "Ten advances in mathematics and theoretical computer science," presenting machine-generated solutions to ten previously unsolved problems spanning pure mathematics and theoretical computer science. The report reads less like a mathematics paper and more like an introduction: OpenAI used it to unveil what the company is internally calling "Astra," described as its next major model family. Rather than launching Astra through a chat interface or a benchmark leaderboard, OpenAI chose to show what an internal, non-public version of the model produced when applied to problems that had resisted proof for years. As of this writing, Astra has not been released publicly and no release date has been announced; what exists is a demonstration of capability, not a product.

Feature Overview

According to OpenAI's report, Astra is designed for long-horizon work: multiple model instances collaborating for hours or days on a single hard problem, planning an approach, revising it as dead ends surface, and delegating sub-problems to other instances running in the background. That framing marks a departure from the single-pass, one-shot query pattern most current models are built around.

The ten results OpenAI published span several distinct fields: high-dimensional geometry, coding theory, group theory, quantum complexity, lattice-based cryptography, and extremal combinatorics. Specific highlights cited in the report include a proof establishing the existence of a non-sofic group, a disproof of Connes's rigidity conjecture, solutions to three separate Erdős problems (numbered 183, 146, and 180), an improved upper bound on sphere-packing density at the Cohn-Elkies threshold, and a new lower bound of roughly n^4/log n for the complexity of computing the permanent of a matrix.

The detail that separates this release from prior AI-and-mathematics claims is that every one of the ten arguments was formalized in Lean, a proof assistant that checks each logical step against a formal kernel rather than relying on natural-language exposition. OpenAI published the corresponding Lean code in a public repository, github.com/openai/ten-proofs, giving outside mathematicians a machine-checkable certificate that each proof's individual steps are logically valid, independent of whether the underlying problem's significance is agreed upon. Human researchers reportedly worked from Astra's formalized arguments to produce conventional written papers for several of the results.

OpenAI also disclosed a cost figure: producing all ten results consumed roughly $2,000 in compute, priced at OpenAI's internal "Sol" API rates. That number is notable mainly as a contrast to the scale of the mathematical problems involved, though it reflects internal pricing rather than a rate available to outside users, since Astra itself isn't public.

Usability Analysis

Because Astra is not available outside OpenAI, there is no way to independently evaluate hands-on usability the way a shipped product would allow. What the report does make clear is the intended interaction model: instead of a single prompt-and-response exchange, Astra is meant to run as a background research process, coordinating several of its own instances over an extended period on one problem, and surfacing partial progress or revised strategies as it goes. That points toward a workflow closer to overseeing an autonomous research collaborator than querying a chatbot, but until Astra ships, that remains a description rather than a demonstrated user experience.

Separately, OpenAI reportedly demonstrated the system to U.S. policymakers and regulators in Washington, and according to reporting from The Decoder, Astra is expected to be the first OpenAI model to go through a new U.S. federal pre-release review framework. Neither OpenAI's own report nor any other primary source independently confirms the details of that framework at this time; it is included here as sourced reporting, not a verified OpenAI claim.

Pros and Cons

Pros

  1. Every proof is formalized in Lean and published on GitHub, giving outside mathematicians a way to independently check the logical validity of each step, rather than having to trust a natural-language claim.
  2. The ten results span six distinct mathematical and computer-science subfields, suggesting the underlying reasoning approach generalizes rather than being tuned to one narrow problem type.
  3. The reported cost, about $2,000 for all ten results, is strikingly low relative to the scale of the problems, at least at OpenAI's internal compute pricing.
  4. Independent mathematician Thomas Bloom described the results as "big news," calling them more significant in terms of new mathematical constructions than earlier AI-assisted math claims.

Cons

  1. Astra itself is not publicly available; there is no benchmark suite, pricing, or release date, so none of this can yet be evaluated as a product.
  2. Lean formalization confirms that each proof's steps are logically valid; it does not, by itself, establish how mathematically significant or novel the results are, and that judgment currently rests on a small number of outside experts rather than broad peer review.
  3. The model family's name, "Astra," is described by OpenAI as tentative; the company has not settled on whether the technology will ship as GPT-6, GPT-5.7, or an entirely separate model class alongside its other internal projects, code-named Sol, Terra, and Luna.
  4. The $2,000 cost figure is self-reported by OpenAI at internal API pricing and has not been independently itemized or verified.

Outlook

If Astra's long-horizon, multi-agent approach holds up under public scrutiny, it would mark a shift in how frontier labs frame model capability: away from single-turn benchmark scores and toward extended, semi-autonomous research work whose output can be independently checked, in this case through formal verification rather than trust in the model's own explanation. The use of Lean to back mathematical claims is itself a meaningful methodological choice; it gives outside reviewers a way to verify correctness without having to take OpenAI's presentation at face value, even if it can't settle questions of significance.

The reported federal pre-release review process is also worth watching, both because it would be a first for OpenAI and because it could set a template other labs are asked to follow for future frontier-class releases. None of this resolves the more basic question of when, or in what form, Astra will actually be available.

Conclusion

This is a capability demonstration, not a product launch. OpenAI showed that an internal, unreleased model could produce ten formally verified solutions to previously unsolved problems in mathematics and theoretical computer science, at a modest reported compute cost, and framed it as an early look at its next model family. Absent a public benchmark suite, a release date, or independent confirmation of the model's broader capabilities, that framing should be read cautiously. It is most useful to researchers and mathematicians tracking formally verified AI reasoning, and to anyone following how frontier labs are beginning to combine extended-horizon agentic work with independently checkable output; it is not yet relevant to anyone evaluating a product they can use.

Editor's Verdict

OpenAI Astra Solves Ten Open Math Problems, Lean-Verified earns a solid recommendation within the gpt space.

The strongest case for paying attention is every proof is formalized in Lean and published on GitHub, allowing independent verification of each proof's logical steps, which raises the bar for what readers should now expect from peers in this space. Reinforcing that, results span six distinct mathematical and computer-science subfields, suggesting generalized rather than narrowly tuned reasoning adds practical value rather than just headline appeal. The broader signal worth registering is straightforward: astra is not publicly released — there is no benchmark suite, pricing, or availability date, only a demonstration via ten formally verified math proofs. On the other side of the ledger, astra is not publicly available — no benchmark suite, pricing, or release date exists yet is a real constraint, not a marketing footnote, and it should factor into any serious decision. Layered on top of that, formal verification confirms logical validity but not mathematical significance, which is judged by a small set of outside experts rather than broad peer review narrows the set of teams for whom this is an obvious yes.

For ChatGPT power users, OpenAI API customers, and enterprise teams already running on the OpenAI stack, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Pros

  • Every proof is formalized in Lean and published on GitHub, allowing independent verification of each proof's logical steps
  • Results span six distinct mathematical and computer-science subfields, suggesting generalized rather than narrowly tuned reasoning
  • Reported cost of roughly $2,000 for all ten results is strikingly low relative to the scale of the problems solved
  • An independent mathematician (Thomas Bloom) called the results more significant, in terms of new mathematical constructions, than prior AI math claims

Cons

  • Astra is not publicly available — no benchmark suite, pricing, or release date exists yet
  • Formal verification confirms logical validity but not mathematical significance, which is judged by a small set of outside experts rather than broad peer review
  • The model family's name and class (GPT-6, GPT-5.7, or a separate line) remain officially undecided
  • The $2,000 cost figure is self-reported at internal pricing and has not been independently itemized

Comments0

Key Features

1. Astra is OpenAI's next major model family, designed for long-horizon, multi-agent work spanning hours or days on a single problem 2. Ten previously unsolved problems solved across six fields: high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics 3. Every proof formalized in Lean and published at github.com/openai/ten-proofs, giving machine-checkable correctness certificates 4. Total compute cost of approximately $2,000 (OpenAI's internal Sol API pricing) to produce all ten results 5. Astra is reportedly the first OpenAI model set to go through a new U.S. federal pre-release review framework, per The Decoder's reporting

Key Insights

  • Astra is not publicly released — there is no benchmark suite, pricing, or availability date, only a demonstration via ten formally verified math proofs.
  • Lean formalization confirms the logical validity of each proof's steps, but it does not by itself establish mathematical significance, which is currently judged by a small number of outside mathematicians such as Thomas Bloom.
  • The results span six distinct subfields — geometry, coding theory, group theory, quantum complexity, cryptography, and combinatorics — pointing to generalized reasoning rather than a single tuned skill.
  • OpenAI reports a total cost of roughly $2,000 to produce all ten results, at internal Sol API pricing, a figure that has not been independently itemized or verified.
  • Astra's design emphasizes long-horizon, multi-agent collaboration — instances that plan, revise, and delegate over hours or days — a different interaction model from single-turn chat.
  • The model family's eventual name and positioning (GPT-6, GPT-5.7, or a separate class alongside Sol, Terra, and Luna) remains undecided according to OpenAI.
  • According to The Decoder's reporting, Astra is expected to be the first OpenAI model to go through a new U.S. federal pre-release review process, following a demonstration to policymakers in Washington.
  • Human researchers reportedly converted Astra's formalized Lean arguments into conventional written mathematical papers for several of the ten results.

Was this review helpful?

Share

Twitter/X