On September 3, 2026, OpenAI introduced GPT‑6 Astra, its new flagship model. OpenAI is calling it the most intelligent and most aligned model it has built, and according to the company, it's now state-of-the-art across computer use, coding, browsing, cybersecurity, science, and general professional work.
That's OpenAI's own framing, worth saying upfront, since everything below comes from their announcement and their benchmarks, not independent testing. Here's what they actually claim, translated out of benchmark-speak.
The headline numbers, in plain English
OpenAI leans heavily on benchmark scores to make its case, and a few of them stand out even without technical context. On ARC-AGI-3, a test built around solving novel puzzle-like environments, Astra scored 99.9%, which OpenAI says is close to matching how efficiently a human would solve the same problems. That's a meaningfully large jump from the roughly 8% the previous model reportedly scored on the same test.
The more interesting claim isn't a leaderboard number at all. OpenAI says Astra helped contribute to two genuinely unsolved math problems, both about the gaps between prime numbers, pushing past bounds that had stood for over a decade in one case and more than 80 years in the other. Benchmark scores are easy to be skeptical of. An actual contribution to open mathematics research is a harder claim to wave away, if it holds up to outside scrutiny.
ChatGPT
ChatGPT is OpenAI's AI assistant for chat, writing, coding, image generation, and everyday tasks, with one of the largest plugin and app ecosystems available.
Visit ChatGPTWhat it's actually meant to be good at
Strip away the benchmark names and the practical pitch is: Astra is meant to be noticeably better at doing real tasks on a computer, not just answering questions about them. OpenAI's examples include filling out online forms, updating records in a CRM, organizing a calendar, researching a topic and drafting a summary, and building and QA-testing a simple website, largely on its own.
OpenAI also claims real speed gains, not just accuracy gains, saying Astra completes computer-use tasks in about 47% less time than the previous model on one internal benchmark, and around 1.9 times faster on another, paired with an updated version of Codex, OpenAI's coding tool.
For document-style work, OpenAI says Astra is better at matching existing templates, producing slides, spreadsheets, and reports that follow a company's existing format and tone rather than generic output that needs reformatting afterward.
Science and health
OpenAI is also pitching Astra as a genuine research tool, not just a chat assistant that happens to know some science. Beyond the prime number contribution mentioned above, OpenAI says Astra set new highs across a range of science and health benchmarks, including a graduate-level science reasoning test (GPQA Diamond, where OpenAI reports a 96% score) and a professional medical knowledge benchmark (HealthBench Professional).
More practically, OpenAI says Astra can work directly inside specialized scientific software, not just talk about the science in the abstract. Their examples include inspecting DNA sequencing data for quality issues and visualizing genetic variation, the kind of hands-on task a researcher would otherwise do manually, to help figure out what's worth investigating further.
Better at reading between the lines
One of the more practically useful claims in the announcement has nothing to do with a benchmark score. OpenAI says Astra is better at handling instructions that leave room for interpretation, filling in routine gaps itself rather than stopping to ask, but asking a focused question when the answer could genuinely change the outcome. In Codex specifically, it can ask that question without blocking, continuing on unrelated work while it waits for a reply, and if nobody answers, it proceeds using a sensible assumption rather than stalling.
OpenAI also says Astra is better at staying oriented as a task evolves partway through. Earlier models, according to OpenAI, sometimes treated a follow-up or correction as a brand new request, losing track of the original goal or constraints that were already agreed. Astra is designed to absorb new instructions without dropping the thread of what it was already doing, a distinction OpenAI backs with a quote from Harvey, a legal AI company, saying Astra approaches legal work "the way a discerning lawyer does," distinguishing established documents from assumptions and flagging gaps rather than guessing past them.
Coding
OpenAI is calling Astra its best model yet for software engineering. The more specific claim worth noting: Codex can now keep working notes across a long coding session instead of repeatedly summarizing and losing detail as the conversation gets longer, something OpenAI says has been a real limitation until now, useful for anyone debugging a gnarly issue or working through a large refactor in one sitting.
The cybersecurity angle
This is the part of the announcement that's genuinely notable, not just marketing framing. OpenAI says Astra crosses into what it calls a "Critical" risk threshold for cyber capability, under its own internal safety framework. During testing, OpenAI says Astra found two real, previously unknown software vulnerabilities, which the company says it's now disclosing responsibly to the maintainers involved.
Because of that, OpenAI says it's intentionally holding back Astra's more advanced offensive cybersecurity capabilities at launch, things like generating working exploits, and plans to loosen those restrictions gradually through a separate access program rather than all at once. It's a rare case of a company publicly admitting a new model is capable of something they're specifically choosing not to fully unlock yet.
Safety and alignment claims
OpenAI says Astra is meaningfully less likely to overstep the boundaries of a task than its previous model. In one internal test, OpenAI says the older model went beyond its authorized scope in 48% of cases when safeguards were removed, compared to 0% for Astra under the same conditions. OpenAI also says Astra is more upfront about what it can and can't actually do, making fewer misleading claims about its own capabilities than its predecessor.
All of this comes from OpenAI's own internal evaluations, so it's their claim to make, not an independently verified result.
Who actually gets it, and when
Astra started rolling out on launch day to a limited set of organizations, with OpenAI saying it would reach all ChatGPT Plus, Pro, Business, and Enterprise users over the following days, alongside availability through the OpenAI API, Microsoft Azure, and Amazon Bedrock. Pro, Business, and Enterprise users also get access to a separate "Astra Pro" variant. For Enterprise accounts specifically, it's off by default, an admin has to switch it on for their workspace.
Pricing
For developers building on the API rather than using ChatGPT directly, OpenAI's standard pricing for Astra is $10 per million input tokens and $50 per million output tokens, with a faster processing option available at double that price. For everyday ChatGPT users, Astra usage is included in existing subscription plans, with the option to buy extra usage credits if you go past your plan's allowance.
