r/MachineLearning
7/29/2026
![AI Security Leaderboard: benchmarking model robustness [P]](https://preview.redd.it/849sof2qs8gh1.jpeg?width=640&crop=smart&auto=webp&s=e35866046cddc5468b5da927b9bf8e5e3eef66c3)
AI Security Leaderboard: benchmarking model robustness [P]
Short summary
A new AI Security Leaderboard ranks frontier models by robustness against 1500 automatically generated jailbreak attempts, measuring universal jailbreaks across domains like offensive cybersecurity and CBRNE. The authors find a significant gap between the most and least robust models and are soliciting community feedback on methodology, including how to fairly compare open-weight versus proprietary models. Planned improvements include adding agent hijacking domains, more realistic agentic tasks, and stronger adaptive optimization attacks.
- •New leaderboard benchmarks frontier model security via 1500 auto-generated jailbreak attempts
- •Significant robustness gap found between most and least secure models
- •Authors seek feedback on open-weight comparison, new domains, and stronger attack methods
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



