< BACK TO NEWS
importantSYS.SOURCE: OpenAI2026-07-08T21:03:51Z

Audit Reveals 30% of SWE-Bench Pro Tasks Contain Evaluation Flaws

An audit of SWE-Bench Pro found ~30% of tasks have critical evaluation flaws, including overly strict tests and underspecified prompts. The study combined automated pipelines and human reviews to identify these issues, highlighting challenges in creating reliable coding benchmarks.

Comments

Read original article

*** END OF TRANSMISSION ***