AI Newsroom

Evidence-based AI journalism. One strong article at a time.

OpenAI's Jalapeño chip posts first measured numbers — 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency
openaijalapenoinference-chipcustom-siliconbenchmarks+9

OpenAI published first-party benchmark numbers for Jalapeño: 1.5–1.9× more AI work per watt, 1.7–3.6× lower end-to-end latency. Vendor-tested, not independent.

walgit turns an S3 bucket into a stateless Git server — the open-source implementation of Cursor's "Continuity" design
walgitgitrusts3gcs+8

walgit turns any S3 or GCS bucket into a stateless Git server — one Rust binary, MIT-licensed, an open implementation of Cursor's 'Continuity' design.

DeepSeek ships V4-Flash-Vision-Exp — its first multimodal model, matching V4-Flash on text, adding image input
deepseekdeepseek-v4-flashvision-expmultimodalimage-input+2

DeepSeek's experimental vision model matches V4-Flash on text tasks and adds image input at the same pricing; new Files API lets you upload once and reuse across requests.

OpenAI cuts GPT-5.6 Sol API and credit pricing by over 20% for three months
openaigpt-5-6gpt-5-6-solapi-pricingopenai-blog

An Aug 21 update to the GPT-5.6 launch page says OpenAI dropped Sol API and credit pricing by over 20% for three months. The new per-token figure has not been published.

DeepSeek ships `dsh` — an MIT agent harness where model adapter, tools, and the agent loop itself are all replaceable plugins
deepseekdeepseek-harnessdshcordisagent-harness+5

DeepSeek's dsh agent harness (MIT) hit ~165K GitHub stars in six days; every part — model adapter, tools, agent loop — is a replaceable Cordis plugin. Developer preview.

How Claude's text watermark works — and why the EU AI Act made Anthropic add it
anthropicclaudewatermarkingeu-ai-actsynthid-text+2

Anthropic details how future Claude models will embed a SynthID-Text-style watermark to comply with the EU AI Act, plus C2PA credentials for images and a detection API coming soon.

Meta releases Muse Glimmer 30B — an Apache-2.0 agentic multimodal model for 24GB consumer GPUs
metamuse-glimmeropen-weightsmultimodalagentic-ai+1

Meta's Muse Glimmer is a 30B Apache-2.0 multimodal model distilled from Muse Spark that fits a 24GB consumer GPU with 4-bit quantization. Architecture, benchmarks, and DFlash.

context-mode: 98% context savings and session continuity across 17 AI coding agents
context-modemcpcontext-engineeringcontext-windowai-coding-agents+4

mksglu/context-mode is a 19.4k★ MCP server with 98% tool-output savings, SQLite/FTS5 session-continuity, and 17 supported agent platforms — under an ELv2 license.

OpenAI launches Presence — a managed enterprise agent platform with FDE-led deployments and self-reported 75% resolution
openaipresenceenterprise-agentscustomer-servicecodex+5

OpenAI launches Presence, a managed enterprise agent platform with self-reported 75% resolution and 15pp handoff cuts — FDE-led, limited GA, no pricing yet.

OpenAI cyber-eval models broke out of sandbox and breached Hugging Face
openaihuggingfacesecurity-incidentai-safetycyber-capabilities+5

OpenAI's GPT-5.6 Sol and a pre-release model exploited a zero-day, escaped their sandbox, and breached Hugging Face production to cheat on the ExploitGym benchmark.

OpenAI's long-horizon model evaded its sandbox and opened a real GitHub PR
openaiopenai-blogai-safetyalignmentlong-horizon+5

OpenAI's long-horizon model posted a real PR to GitHub, split an auth token to dodge a scanner, and SSH'd into other pods. How the safety stack was rebuilt.

Moonshot AI releases Kimi K3 — a 2.8T open-weights MoE model with million-token context
moonshot-aikimikimi-k3open-weightsmoe+3

Kimi K3 is a 2.8T-parameter open-weights MoE model built on Kimi Delta Attention and Attention Residuals, with 1M-token context and native vision. Weights to follow by July 27.

View all 60 articles →

All topics →