Back to feed
r/MachineLearning
r/MachineLearning
6/24/2026
DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]

DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]

Short summary

DeepSWE is an open-source benchmark for evaluating code generation in frontier models, featuring 91 repositories across 5 languages with contamination-free tasks and real-world complexity. Solutions require 5.5x more code than SWE-bench Pro, with hand-written verifiers that test behavior rather than implementation details.

  • Open-source benchmark for code generation models with 91 repos across 5 languages
  • Solves contamination by using original tasks, not adapted from existing commits
  • Solutions require 5.5x more code and 2x more tokens than SWE-bench Pro

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more