Dev.to
7/17/2026

OpenAI audit finds ~30% of SWE-Bench Pro tasks broken; proposes uncertainty budgets for coding benchmarks
Original: If 30% of Coding Tasks May Be Broken, Your Leaderboard Needs an Uncertainty Budget
Short summary
OpenAI's audit of SWE-Bench Pro estimates ~30% of tasks are broken, invalidating the assumption that every benchmark task is a valid trial. The article proposes versioning task validity, preserving disputed cases, and publishing sensitivity bounds showing how conclusions change across plausible denominators. It includes a Python calculator for validity-sensitivity intervals and invariants for append-only attempts, deterministic recomputation, and common-task-set pairwise comparisons.
- •~30% of SWE-Bench Pro tasks are broken per OpenAI's audit, making raw leaderboard rankings unreliable
- •Proposes a validity taxonomy (unreviewed/valid/broken/disputed) and sensitivity bounds instead of confidence intervals
- •Includes a Python calculator and invariants for provenance, append-only attempts, and common-task-set comparisons
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



