Back to feed
Dev.to
Dev.to
7/10/2026
The original title is: "OpenAI just found ~30% of SWE-Bench Pro is broken — and retracted their own recommendation"

The original title is: "OpenAI just found ~30% of SWE-Bench Pro is broken — and retracted their own recommendation"

Original: OpenAI just found ~30% of SWE-Bench Pro is broken — and retracted their own recommendation

Short summary

OpenAI audited SWE-Bench Pro and found ~30% of its 731 tasks are broken due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. They retracted their earlier endorsement, noting that benchmarks built from GitHub PRs don't produce clean model-evaluation tasks. OpenAI used Codex-based investigator agents to run the audit and is calling for new benchmarks designed specifically to test model capabilities.

  • ~30% of SWE-Bench Pro tasks are broken; OpenAI retracted its recommendation
  • Four failure patterns: strict tests, underspecified prompts, low-coverage tests, misleading prompts
  • OpenAI calls for new benchmarks built specifically for model evaluation, not repurposed PRs

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more