Dev.to
7/10/2026

The original title is: "OpenAI just found ~30% of SWE-Bench Pro is broken — and retracted their own recommendation"
Original: OpenAI just found ~30% of SWE-Bench Pro is broken — and retracted their own recommendation
Short summary
OpenAI audited SWE-Bench Pro and found ~30% of its 731 tasks are broken due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. They retracted their earlier endorsement, noting that benchmarks built from GitHub PRs don't produce clean model-evaluation tasks. OpenAI used Codex-based investigator agents to run the audit and is calling for new benchmarks designed specifically to test model capabilities.
- •~30% of SWE-Bench Pro tasks are broken; OpenAI retracted its recommendation
- •Four failure patterns: strict tests, underspecified prompts, low-coverage tests, misleading prompts
- •OpenAI calls for new benchmarks built specifically for model evaluation, not repurposed PRs
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



