Back to feed
MarkTechPost
MarkTechPost
7/16/2026
OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection

OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection

Short summary

OpenAI trained GPT-Red, an internal attacker model using self-play RL against defender LLMs, which outperformed human red-teamers 84% to 13% on indirect prompt injection. The model discovered a novel 'Fake Chain-of-Thought' attack class and reduced GPT-5.6 Sol's direct injection failures by 6x. OpenAI acknowledges remaining weaknesses in multi-turn and image-based attack scenarios.

  • GPT-Red beats human red-teamers 84% to 13% on prompt injection
  • Discovered novel Fake Chain-of-Thought attack class
  • Cut GPT-5.6 Sol injection failures 6x; still struggles with multi-turn and image attacks

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more