Cybersecurity
OpenAI Unveils 'GPT-Red' for Automated AI Model Security Testing
OpenAI on 15 July 2026 unveiled GPT-Red, an automated red-teaming system designed to identify vulnerabilities in AI models before they are deployed.
The company said the system was trained through self-play reinforcement learning and achieved an 84% attack success rate on novel scenarios, compared with just 13% for human red-teamers, successfully breaking several advanced models.
OpenAI added that integrating GPT-Red into the training of its latest model reduced the model's failures against prompt-injection attacks sixfold, as part of what it described as a "self-improving safety flywheel" in which current models help make future models more robust.
Source: OpenAI (openai.com) — 15 July 2026
📲Follow AI news on WhatsAppVerified news, straight from official sources