OpenAI retracts SWE-Bench Pro: ~30% of tasks broken, audit finds
OpenAI's own datapoint + 5-engineer audit flags 27–34% of SWE-Bench Pro tasks as broken. It just retracted its own recommendation to use the benchmark.
CodeBurn: free, local-first cost tracker for 31 AI coding tools
CodeBurn (getagentseal/codeburn) reads the session files your AI coding tools already write, breaks down every token and dollar across 31 integrations, MIT, no proxy, no API key.