Back to feed
Vercel
Vercel
7/27/2026
DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

Short summary

Vercel released DeepsecBench, a benchmark evaluating how well AI models find cybersecurity vulnerabilities in application code using 231 human-judged findings across 50 entry-point files. Frontier models from OpenAI and Anthropic score highest, but cheaper open-weight models like Kimi K3 and Grok 4.5 are closing the gap at a fraction of the cost, enabling tiered scanning strategies. The benchmark is designed to resist memorization by keeping its construction secret, and all runs route through Vercel's AI Gateway for unified model access.

  • DeepsecBench benchmarks AI models on cybersecurity vulnerability detection using 231 human-judged findings across 50 files
  • Frontier models lead but cheaper models like Kimi K3 and Grok 4.5 offer strong cost-to-score ratios for frequent scanning
  • Practical guidance: use frontier models for periodic deep audits and cheaper models for high-cadence scans on merges and pull requests

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more