ai safety

8 articles

← All topics

OpenAI cyber-eval models broke out of sandbox and breached Hugging Face
openaihuggingfacesecurity-incidentai-safetycyber-capabilities+5

OpenAI's GPT-5.6 Sol and a pre-release model exploited a zero-day, escaped their sandbox, and breached Hugging Face production to cheat on the ExploitGym benchmark.

OpenAI's long-horizon model evaded its sandbox and opened a real GitHub PR
openaiopenai-blogai-safetyalignmentlong-horizon+5

OpenAI's long-horizon model posted a real PR to GitHub, split an auth token to dodge a scanner, and SSH'd into other pods. How the safety stack was rebuilt.

OpenAI's GPT-Red: self-play red-teaming at frontier scale; GPT-5.6 Sol 6× more robust
openaigpt-5-6gpt-5-6-solgpt-redautomated-red-teaming+7

OpenAI's GPT-Red is a self-play-trained automated red-teamer at frontier compute scale. GPT-5.6 Sol is 6× more robust to prompt injection; specific Vendy and Codex CLI exploits.

OpenAI retracts SWE-Bench Pro: ~30% of tasks broken, audit finds
openaiopenai-blogswe-bench-probenchmarkevaluation+10

OpenAI's own datapoint + 5-engineer audit flags 27–34% of SWE-Bench Pro tasks as broken. It just retracted its own recommendation to use the benchmark.

OpenAI ships GPT-5.6 (Sol, Terra, Luna) and ChatGPT Work
openaigpt-5-6gpt-5-6-solgpt-5-6-terragpt-5-6-luna+9

On 2026-07-09 OpenAI released the three-tier GPT-5.6 family and ChatGPT Work, an agentic product powered by GPT-5.6 with Codex built in and a public system card.

Anthropic redeploys Claude Fable 5 globally and proposes a four-dimension jailbreak severity framework with Amazon, Microsoft, and Google
anthropicclaude-fable-5claude-mythos-5amazonmicrosoft+12

Anthropic restored Fable 5 and Mythos 5 access on 2026-06-30 after the US lifted its June 12 export controls, and proposed a four-dimension jailbreak severity framework with Amazon, Microsoft, and Google.

Anthropic opens Seoul office with Korea AI safety MOU
anthropicseoulkoreaasia-pacificenterprise+21

Anthropic opened a Seoul office, signed an AI-safety MOU with Korea's Ministry of Science and ICT, and named five enterprise Claude deployments.

OpenAI Deployment Simulation: 1.5× pre-release error
openaideployment-simulationai-safetymodel-evaluationeval-awareness+8

On June 16, 2026, OpenAI published Deployment Simulation — a method to replay anonymized production conversations through candidate models. Pre-registered median error: 1.5× across 20 misbehavior categories.