OpenAI retracts SWE-Bench Pro: ~30% of tasks broken, audit finds
OpenAI's own datapoint + 5-engineer audit flags 27–34% of SWE-Bench Pro tasks as broken. It just retracted its own recommendation to use the benchmark.
Sonnet 5 launches at $2/$10, nears Opus 4.8
Anthropic's Claude Sonnet 5 launches 2026-06-30 with intro pricing $2/$10 per MTok (→$3/$15 on Sep 1), a new effort-level API dial, and a near-Opus cost-performance curve.