MIT Technology Review
7/15/2026

OpenAI's GPT-Red: an automated red-teaming LLM for hardening model cybersecurity defenses
Original: Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
Short summary
OpenAI has developed GPT-Red, an automated LLM red-teaming system that acts as a sparring partner to harden its models against cyberattacks. Training GPT-5.6 against GPT-Red reportedly produced OpenAI's most robust model release to date. The tool automates adversarial testing to find and patch vulnerabilities before deployment.
- •OpenAI built GPT-Red, an automated LLM red-teaming system for cybersecurity
- •GPT-5.6 was trained against GPT-Red, making it OpenAI's most robust release
- •The approach automates adversarial testing to improve model defenses
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



