r/MachineLearning
6/24/2026
![DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]](https://preview.redd.it/lacvagyr159h1.png?width=140&height=89&auto=webp&s=14f97a97511fbfe2fd767e4dc986ce0b4da5c73e)
DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]
Short summary
DeepSWE is an open-source benchmark for evaluating code generation in frontier models, featuring 91 repositories across 5 languages with contamination-free tasks and real-world complexity. Solutions require 5.5x more code than SWE-bench Pro, with hand-written verifiers that test behavior rather than implementation details.
- •Open-source benchmark for code generation models with 91 repos across 5 languages
- •Solves contamination by using original tasks, not adapted from existing commits
- •Solutions require 5.5x more code and 2x more tokens than SWE-bench Pro
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



