preparedness framework

3 articles

← All topics

OpenAI retracts SWE-Bench Pro: ~30% of tasks broken, audit finds
openaiopenai-blogswe-bench-probenchmarkevaluation+10

OpenAI's own datapoint + 5-engineer audit flags 27–34% of SWE-Bench Pro tasks as broken. It just retracted its own recommendation to use the benchmark.

OpenAI ships GPT-5.6 (Sol, Terra, Luna) and ChatGPT Work
openaigpt-5-6gpt-5-6-solgpt-5-6-terragpt-5-6-luna+9

On 2026-07-09 OpenAI released the three-tier GPT-5.6 family and ChatGPT Work, an agentic product powered by GPT-5.6 with Codex built in and a public system card.

OpenAI ships GPT-Live — a full-duplex voice model that listens and speaks at the same time
openaigpt-livevoice-modelsfull-duplexchatgpt+5

OpenAI's GPT-Live-1 and GPT-Live-1 mini are full-duplex voice models that can listen and speak at once, delegate to GPT-5.5, and ship as the new ChatGPT default from 2026-07-08.