swe bench pro

2 articles

← All topics

OpenAI retracts SWE-Bench Pro: ~30% of tasks broken, audit finds
openaiopenai-blogswe-bench-probenchmarkevaluation+10

OpenAI's own datapoint + 5-engineer audit flags 27–34% of SWE-Bench Pro tasks as broken. It just retracted its own recommendation to use the benchmark.

Sonnet 5 launches at $2/$10, nears Opus 4.8
anthropicclaude-sonnet-5claude-opus-4-8anthropic-pricinganthropic-api+20

Anthropic's Claude Sonnet 5 launches 2026-06-30 with intro pricing $2/$10 per MTok (→$3/$15 on Sep 1), a new effort-level API dial, and a near-Opus cost-performance curve.