high risk claim

10 articles

← All topics

OpenAI launches Presence — a managed enterprise agent platform with FDE-led deployments and self-reported 75% resolution
openaipresenceenterprise-agentscustomer-servicecodex+5

OpenAI launches Presence, a managed enterprise agent platform with self-reported 75% resolution and 15pp handoff cuts — FDE-led, limited GA, no pricing yet.

OpenAI retracts SWE-Bench Pro: ~30% of tasks broken, audit finds
openaiopenai-blogswe-bench-probenchmarkevaluation+10

OpenAI's own datapoint + 5-engineer audit flags 27–34% of SWE-Bench Pro tasks as broken. It just retracted its own recommendation to use the benchmark.

OpenAI ships GPT-5.6 (Sol, Terra, Luna) and ChatGPT Work
openaigpt-5-6gpt-5-6-solgpt-5-6-terragpt-5-6-luna+9

On 2026-07-09 OpenAI released the three-tier GPT-5.6 family and ChatGPT Work, an agentic product powered by GPT-5.6 with Codex built in and a public system card.

Anthropic redeploys Claude Fable 5 globally and proposes a four-dimension jailbreak severity framework with Amazon, Microsoft, and Google
anthropicclaude-fable-5claude-mythos-5amazonmicrosoft+12

Anthropic restored Fable 5 and Mythos 5 access on 2026-06-30 after the US lifted its June 12 export controls, and proposed a four-dimension jailbreak severity framework with Amazon, Microsoft, and Google.

DeepSeek releases DeepSpec: open-source full-stack for speculative decoding
deepseekdeepspecdflashdsparkeagle3+12

DeepSeek published DeepSpec, a full-stack MIT-licensed codebase for training and evaluating draft models for speculative decoding, bundling DSpark, DFlash, and Eagle3.

US government is now a customer gatekeeper for OpenAI Sol and Claude Mythos 5
openaianthropicgpt-5-6-solclaude-mythos-5claude-fable-5+14

On 2026-06-26 OpenAI began a GPT-5.6 Sol preview with the customer list coordinated with the US government; the same day, Commerce lifted the export block on Anthropic's Claude Mythos 5.

oh-my-pi: a terminal agent that treats the harness as the product
oh-my-piompai-coding-agentharnessharness-engineering+25

can1357/oh-my-pi: a Rust-based terminal AI coding agent (fork of Mario Zechner's pi-mono) with Hashline edits, an advisor model, 40+ providers, 18-hour release cadence.

caveman: terse-output skill, 75k stars in 11 weeks
open-sourcetypescriptclaude-codecodexgemini+21

GitHub JuliusBrussee/caveman: a TypeScript skill for 30+ agent platforms that asks the agent to drop filler. 75k stars, MIT, 65% output-token reduction.

OpenAI ships ChatGPT health; o3 re-solves 4.8% of rare
openaichatgptgpt-5-5-instanthealth-airare-disease+11

OpenAI's June 18, 2026 release: GPT-5.5 Instant rated higher than physicians on a 3,500-response panel; o3 Deep Research surfaces 18 of 376 unsolved rare-disease cases (4.8%).

DiffusionGemma: 1,000+ tokens/sec open-weights text gen
google-deepmindgemmadiffusiongemmatext-diffusionopen-weights+10

Google DeepMind's June 10, 2026 release: a 26B/3.8B text-diffusion model denoising 256 tokens in parallel; ~4x faster than AR on a single H100, Apache 2.0, explicitly experimental.