OpenAI's upcoming AI model Astra has crossed its highest cybersecurity capability threshold, becoming the first model the company believes can autonomously discover previously unknown vulnerabilities and develop ways to exploit them across hardened computer systems.
The company has classified Astra at the “Critical” cybersecurity capability level under its Preparedness Framework, a step above the “High” designation given to GPT-5.6 Sol and its other recent frontier models.
The designation means that, when provided with the necessary tools and system access, Astra is capable of identifying previously unknown security flaws and developing functional exploits across well-protected systems without requiring a human operator to guide each stage of the process.
Interestingly, Astra is the first model it has classified at this level.
AI finds zero-day vulnerabilities
Testing points to a substantial increase in autonomous cybersecurity capabilities.
Astra achieved a 100% score on ExploitBench, which measures the ability to develop exploits from known vulnerabilities.
OpenAI subsequently tested the model against an internal benchmark containing 20 recently disclosed high-severity vulnerabilities. During those evaluations, Astra discovered and used two previously unknown zero-day vulnerabilities as part of an exploit chain. OpenAI said it was working to disclose the vulnerabilities to their maintainers.
Expert testing also placed Astra against hardened browser and operating-system environments.
According to OpenAI, the model discovered previously unknown browser vulnerabilities and combined them into an exploit chain capable of escaping a browser sandbox and executing commands on the host computer.
It separately identified multiple vulnerabilities in a hardened operating system and combined them into a privilege-escalation chain capable of moving from an unprivileged account to root access.
Stronger safeguards before release
Those capabilities have also forced OpenAI to strengthen the safeguards surrounding Astra.
The company delayed parts of the model's development and release while additional protections against cyber misuse and unauthorised model actions were implemented and tested.
OpenAI said it had also paused some frontier training following a separate security incident involving Hugging Face, before restarting a major reinforcement-learning run on August 28 after new security requirements were put in place.
Protections for Astra include stronger refusal training, system-level safeguards, monitoring designed to detect potentially unauthorised activity and additional restrictions around access to its most powerful cybersecurity capabilities.
Access to remain restricted
OpenAI plans to make Astra available soon, although its most advanced cyber capabilities will not initially be opened broadly.
Advanced cybersecurity functionality will first be made available to selected testers, with expanded defensive access expected through OpenAI's Daybreak Blue program.
For everyday users, the safeguards could occasionally result in legitimate tasks being slowed, paused or stopped if automated monitoring detects activity that resembles malicious or unauthorised cybersecurity work.
The development marks an important transition for generative AI. The cybersecurity debate is no longer simply about whether AI can help programmers identify vulnerabilities.
With Astra, OpenAI believes a frontier model has reached the point where it can independently discover weaknesses, develop exploits and assemble complex attack chains — raising both the potential value of AI for cyber defenders and the consequences if equivalent capabilities are misused.