Back to feed
Dev.to
Dev.to
7/17/2026
OpenAI audit finds ~30% of SWE-Bench Pro tasks broken; proposes uncertainty budgets for coding benchmarks

OpenAI audit finds ~30% of SWE-Bench Pro tasks broken; proposes uncertainty budgets for coding benchmarks

Original: If 30% of Coding Tasks May Be Broken, Your Leaderboard Needs an Uncertainty Budget

Short summary

OpenAI's audit of SWE-Bench Pro estimates ~30% of tasks are broken, invalidating the assumption that every benchmark task is a valid trial. The article proposes versioning task validity, preserving disputed cases, and publishing sensitivity bounds showing how conclusions change across plausible denominators. It includes a Python calculator for validity-sensitivity intervals and invariants for append-only attempts, deterministic recomputation, and common-task-set pairwise comparisons.

  • ~30% of SWE-Bench Pro tasks are broken per OpenAI's audit, making raw leaderboard rankings unreliable
  • Proposes a validity taxonomy (unreviewed/valid/broken/disputed) and sensitivity bounds instead of confidence intervals
  • Includes a Python calculator and invariants for provenance, append-only attempts, and common-task-set comparisons

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more