MarkTechPost
7/16/2026

OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection
Short summary
OpenAI trained GPT-Red, an internal attacker model using self-play RL against defender LLMs, which outperformed human red-teamers 84% to 13% on indirect prompt injection. The model discovered a novel 'Fake Chain-of-Thought' attack class and reduced GPT-5.6 Sol's direct injection failures by 6x. OpenAI acknowledges remaining weaknesses in multi-turn and image-based attack scenarios.
- •GPT-Red beats human red-teamers 84% to 13% on prompt injection
- •Discovered novel Fake Chain-of-Thought attack class
- •Cut GPT-5.6 Sol injection failures 6x; still struggles with multi-turn and image attacks
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



