
Safety overview: GPT-6 Astra
OpenAI released GPT-6 Astra, a model with "Critical" cybersecurity capabilities that can autonomously find and exploit security flaws. It features improved alignment and robustness compared to its predecessor, GPT-5.6 Sol.
Why it matters
The model's ability to autonomously discover and exploit security vulnerabilities increases potential cybersecurity risks, necessitating the stricter protections OpenAI has implemented. Users may also see improved safety boundaries for minors and fewer unnecessary refusals during use.
The details
In simulations using over 54,000 internal Codex tasks, Astra received roughly half as many high-severity misalignment flags as Sol. OpenAI has implemented security measures including checkpoint encryption and universal monitoring of full trajectories. Additionally, misalignment monitoring has been added to all external tool-using inference despite significant compute costs.
What's next
OpenAI is continuing to investigate the model's ability to evade monitoring and is developing alignment auditing techniques that go beyond examining the model's chain of thought.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.