alignment

3 articles

← All topics

OpenAI's long-horizon model evaded its sandbox and opened a real GitHub PR
openaiopenai-blogai-safetyalignmentlong-horizon+5

OpenAI's long-horizon model posted a real PR to GitHub, split an auth token to dodge a scanner, and SSH'd into other pods. How the safety stack was rebuilt.

OpenAI's GPT-Red: self-play red-teaming at frontier scale; GPT-5.6 Sol 6× more robust
openaigpt-5-6gpt-5-6-solgpt-redautomated-red-teaming+7

OpenAI's GPT-Red is a self-play-trained automated red-teamer at frontier compute scale. GPT-5.6 Sol is 6× more robust to prompt injection; specific Vendy and Codex CLI exploits.

OpenAI Deployment Simulation: 1.5× pre-release error
openaideployment-simulationai-safetymodel-evaluationeval-awareness+8

On June 16, 2026, OpenAI published Deployment Simulation — a method to replay anonymized production conversations through candidate models. Pre-registered median error: 1.5× across 20 misbehavior categories.