Better Models, Worse Tools: Claude tool calls regress on Sonnet 5 and Opus 4.8
Armin Ronacher's controlled tests show Claude Opus 4.8 and Sonnet 5 produce ~20% malformed tool calls against Pi's edit tool; older models are clean.
OpenAI Deployment Simulation: 1.5× pre-release error
On June 16, 2026, OpenAI published Deployment Simulation — a method to replay anonymized production conversations through candidate models. Pre-registered median error: 1.5× across 20 misbehavior categories.