
Path to Astra: critical capabilities and frontier safeguards
OpenAI has designated its Astra model as meeting the 'Critical' cybersecurity capability threshold because it can autonomously find and exploit unknown security flaws. This classification requires stronger safeguards during development and before its release.
Why it matters
The ability to automate complex cyberattacks increases security risks, while new tools could help defenders fix vulnerabilities faster. Users may experience more frequent interruptions to AI tasks due to heightened security monitoring.
The details
Astra discovered two zero-day vulnerabilities in an internal benchmark from June–August 2026. The model successfully created a browser-compromise chain that escaped the sandbox and executed commands on the host. It also demonstrated a local privilege-escalation chain from an unprivileged user to root. To prevent misuse, OpenAI uses a stack of post-trained refusals, system classifiers, and chain-of-thought monitoring.
What's next
OpenAI will release a system card at launch containing details on safety, security, and alignment testing. Access to advanced cybersecurity capabilities will expand through Daybreak Blue following the initial alpha testing phase.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.