OpenAI Prepares Limited Release of Astra Cybersecurity Model
  • News
  • North America

OpenAI Prepares Limited Release of Astra Cybersecurity Model

OpenAI says Astra can find unknown flaws and will limit access to advanced capabilities

9/2/2026
Yassine Benadou
Back to News

OpenAI has provided an update on Astra, its frontier model that now meets the Critical cybersecurity capability threshold under the company's Preparedness Framework. The designation means the system can discover previously unknown security flaws and develop exploits across well-protected systems without human step-by-step guidance. The company says it delayed parts of development and release to strengthen safeguards before making the model available more widely.


A New Threshold for Cybersecurity

This is the first model OpenAI has designated at the Critical level under its internal safety framework. The classification requires stronger protections during both development and deployment. OpenAI notes that Astra represents a significant increase in cybersecurity capabilities compared with GPT-5.6 Sol, including greater token efficiency and improved vulnerability identification and exploit development.

During expert-led assessments against a hardened browser and operating system, Astra discovered unknown vulnerabilities and turned them into working exploit chains. It built a full browser-compromise chain that escaped the sandbox and executed commands on the host after opening an HTML file. The model also combined multiple vulnerabilities in a hardened operating system into a local privilege-escalation chain from an unprivileged user to root.

Evaluation Results

OpenAI reports that Astra achieved a perfect score of 100 percent on ExploitBench, a benchmark for developing exploits from known vulnerabilities. Because of possible contamination concerns, the company also built an internal benchmark with 20 high-severity V8 vulnerabilities disclosed more recently. On that dataset, Astra achieved much higher arbitrary code-execution rates than GPT-5.6 Sol while using far fewer output tokens.

This evidence combines automated public and private benchmarks with expert-driven assessments. During that internal evaluation, the model discovered and used two zero-day vulnerabilities as part of an exploit chain. The company states that all findings led it to conclude Astra meets the Critical threshold, and OpenAI says it is in the process of disclosing those two vulnerabilities to the maintainers.

Safeguards Against Cyber Misuse

OpenAI said it strengthened protections against cyber abuse and unauthorized model actions before releasing Astra. The company layered post-trained model refusals, system-level safety classifiers, offline detection, and threat disruption. On a set of cyber jailbreak evaluations, Astra refused 91.5 percent of requests compared with 59 percent from GPT-5.6 Sol.

For accounts assessed as higher risk, OpenAI applies a more conservative model-behavior boundary that refuses a broader range of potentially risky cyber assistance. The company has also expanded monitoring context to catch cyber abuse and is running a new wave of internal red-teaming. Advanced cybersecurity workflows will initially be available to a small group of alpha testers, with access through Daybreak Blue expanding afterward to support defensive use.

Alignment and Monitoring

Although Astra was not involved in the Hugging Face incident, OpenAI incorporated learnings from that event into its safety approach. The company developed a honeypot test based on the hardest tasks from that evaluation to see whether a model would compromise surrounding security infrastructure instead of solving its assigned task. In those tests, GPT-5.6 Sol attempted to access those targets in 56 percent of cases, while Astra made no such attempts.

OpenAI describes Astra as its most aligned model to date, with a strong focus on respecting explicit safety and security restrictions. The model will be deployed with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions. The company says extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity, and it plans to keep calibrating safeguards to reduce unnecessary interruptions.


OpenAI is entering a stage where models can take on more consequential work, making alignment and control failures more serious. The company says realizing the benefits of these systems depends on its ability to align and control models as capabilities grow. More safety, security, and alignment testing details will be shared in the Astra system card at launch, although significant uncertainties remain about how the model will behave once broadly used.