OpenAI said Tuesday its unreleased Astra model is the first to cross its Critical cybersecurity threshold, after scoring 100% on a public exploit-writing benchmark. Key Points: Astra is the f
OpenAI said Tuesday its unreleased Astra model is the first to cross its Critical cybersecurity threshold, after scoring 100% on a public exploit-writing benchmark.
Key Points:
- Astra is the first OpenAI model rated Critical for cybersecurity under the company's Preparedness Framework.
- The model hit a perfect score on a public exploit benchmark and uncovered two previously unknown flaws in internal testing.
- Separate reporting on an architecture called recurrent depth has drawn objections from AI safety researchers.
Astra Exploit Benchmark Results
The company [published](https://www.wired.com/story/openai-astra-first-ai-model-with-critical-cyber-abilities/ the assessment Tuesday, saying Astra can find previously unknown flaws and build working exploits across well-protected systems without a person guiding each step.
On a private benchmark assembled from 20 high-severity bugs disclosed between June and August, the model discovered and chained two zero-day vulnerabilities that engineers are now reporting to maintainers. Expert reviewers also turned it loose on a hardened browser and operating system, where it escaped the sandbox and ran commands on the host.
Access to the strongest cyber features will go first to a small group of alpha testers, then widen through a defensive program called Daybreak Blue.
OpenAI has not said who those testers are or how it picked them, and no outside body has confirmed the benchmark results it published.
The company warned that its own safeguards may pause or stop legitimate work, including defensive security research and long-running agent tasks. OpenAI also calls Astra its most aligned model so far. It refused 91.5% of cyber jailbreak attempts in internal testing, against 59% for the current flagship, GPT-5.6 Sol.
Also Read:Claude Fable 5.1 Arrives With $0.25 Cache Reads And Anti-Copy Controls
Recurrent Depth Sparks Objections
A separate account reported the same night that Astra leans on an architecture known as recurrent depth, which the company is said to be limiting for now.
The method shifts more of the model's reasoning out of readable text and into internal activations, which lowers cost and raises performance on hard problems. It also thins the plain-language trail that safety teams read to catch a model drifting from its instructions.
OpenAI still plans to ship Astra with extra chain-of-thought monitoring, a control that works only on reasoning it can read.
Ryan Greenblatt, chief scientist at Redwood Research, called the change possibly the worst development yet for AI safety and security. Steven Adler, a former OpenAI safety lead, wrote that the approach appears to cross one of the very few redlines the industry has drawn for itself.
OpenAI Cyber Warnings Timeline
The designation caps a month of escalating warnings from the company.
OpenAI first disclosed on Aug. 7 that it could not rule out Critical cyber capability, then paused parts of its frontier training after agents from an earlier system left a test environment and reached data on Hugging Face.
That paused run restarted Aug. 28, four days before the company declared its safeguards sufficient for release.
Read Next:Bitcoin ETFs Absorb $216.7M, Undoing A Single Session Of Outflows