
GPT-Red: Unlocking Self-Improvement for Robustness | OpenAI
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
Show entities and relationshipsHide entities and relationships
In this article
Key connections
OpenAI developed and owns GPT-5.3.
OpenAI develops the GPT-5.6 model family.
OpenAI developed and owns GPT-5.1.
OpenAI owns GPT-5.6 Sol
OpenAI developed GPT-5.6 Sol
Andon Labs owns Vendy
Andon Labs developed and owns Vendy.
Andon Labs owns Project Vend
Andon Labs developed Project Vend.
Show 14 more connectionsShow fewer connections
GPT-Red is built with Reinforcement Learning
GPT-Red is trained using self-play reinforcement learning.
GPT-Red was used to adversarially train GPT-5.6.
GPT-Red is related to GPT-5.6 Sol
GPT-Red was used to adversarially train GPT-5.6 Sol.
GPT-Red was pitted against Vendy in a simulated and live red-teaming exercise.
GPT-Red was used to attack a Codex CLI agent.
GPT-Red was tested against GPT-5.1 on indirect prompt injection scenarios.
GPT-Red was compared to a prompted GPT-5.5 baseline in attacks against Codex.
Vendy is related to Project Vend
Vendy is an autonomous vending machine agent similar to Project Vend.
Codex is related to GPT-5.4 mini
The Codex CLI agent was based on GPT-5.4 mini.
OpenAI uses AI agents to improve next-generation models.
Andon Labs uses AI Agents
Andon Labs builds AI-powered autonomous agent systems.
Andon Labs uses Artificial Intelligence
Andon Labs develops artificial intelligence systems.
OpenAI owns GPT-5.4 mini
OpenAI developed GPT-5.4 mini.
GPT-Red uses Prompt Injection
GPT-Red was developed to test and improve model robustness against prompt injection attacks.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.