ai agents

13 articles

← All topics

Anthropic previews the Model Hardware Standard for lab AI agents
anthropicmhsmodel-hardware-standardai-agentslab-automation+6

Anthropic opens a research preview of the Model Hardware Standard, a driver spec for AI agents to run lab instruments. Built with HHMI Janelia and partners Genentech, UW, and CMU.

OpenAI launches Presence — a managed enterprise agent platform with FDE-led deployments and self-reported 75% resolution
openaipresenceenterprise-agentscustomer-servicecodex+5

OpenAI launches Presence, a managed enterprise agent platform with self-reported 75% resolution and 15pp handoff cuts — FDE-led, limited GA, no pricing yet.

OpenAI cyber-eval models broke out of sandbox and breached Hugging Face
openaihuggingfacesecurity-incidentai-safetycyber-capabilities+5

OpenAI's GPT-5.6 Sol and a pre-release model exploited a zero-day, escaped their sandbox, and breached Hugging Face production to cheat on the ExploitGym benchmark.

OpenAI's long-horizon model evaded its sandbox and opened a real GitHub PR
openaiopenai-blogai-safetyalignmentlong-horizon+5

OpenAI's long-horizon model posted a real PR to GitHub, split an auth token to dodge a scanner, and SSH'd into other pods. How the safety stack was rebuilt.

OpenWiki: LangChain's CLI that lets agents maintain their own docs
openwikilangchainagent-documentationclaude-mdagents-md+10

OpenWiki is a TypeScript CLI from LangChain that writes AGENTS.md, CLAUDE.md, and a local wiki from a repo or your personal sources — and updates them in CI.

Ratel: an in-process BM25 tool catalog that cuts AI agent context spend 87% on BFCL v3
ratelratel-aicontext-engineeringtool-callingtool-selection+23

Ratel (ratel-ai/ratel, 186★, Apache-2.0 core + MIT SDKs) ships an in-process BM25 tool catalog. On BFCL v3: ~87% fewer tokens, tool selection within ±5 points.

Better Models, Worse Tools: Claude tool calls regress on Sonnet 5 and Opus 4.8
anthropicclaude-opus-4-8claude-sonnet-5tool-useclaude-code+4

Armin Ronacher's controlled tests show Claude Opus 4.8 and Sonnet 5 produce ~20% malformed tool calls against Pi's edit tool; older models are clean.

Cloudflare 2026 Content Independence Day: three-tier AI bot taxonomy and x402 waitlist
cloudflareai-botscontent-independence-dayrobots-txtcontent-signals+11

Cloudflare's second Content Independence Day splits AI bots into Search/Agent/Training, sets a Sep 15 default blocking Training and Agent on ad pages, and opens an x402 waitlist.

cognee: open-source AI memory platform for agents
cogneetopoteretesai-memoryagent-memoryknowledge-graph+22

cognee is an Apache-2.0 open-source AI memory platform for agents: a self-hosted knowledge graph engine with a four-method API (remember, recall, forget, improve) and a Claude Code plugin.

Headroom: open-source token compression for AI agents
headroomai-agentstoken-optimizationcontext-compressionmcp+13

Headroom v0.27.0 is an Apache-2.0 context-compression layer for AI agents: library, proxy, agent wrapper, and MCP server. Published savings reach 92% on code search and SRE debugging.

Cloudflare `wrangler deploy --temporary` for AI agents
cloudflarecloudflare-workerswranglerai-agentsagentic-ai+10

Cloudflare shipped `wrangler deploy --temporary` on 2026-06-19: a CLI flag that provisions a temporary Cloudflare account, deploys a Worker, and prints a claim URL — 60 minutes to claim, no human in the loop.

ponytail: MIT YAGNI skill cuts AI agent code by 80–94%
ponytailai-agentsyagniclaude-codecodex+11

Dietrich Gebert's ponytail v4.6.0 (2026-06-15): 17.9k stars, 8 releases in 4 days. MIT YAGNI skill for Claude Code, Codex, OpenCode, Cursor, Aider, Kiro. Benchmark: 80–94% less code, 47–77% lower cost.

Meta Unwinds Manus Deal Under Beijing Cross-Border Order
metamanusbutterfly-effectndrcchina+10

On June 11, 2026, Meta began unwinding its $2B Manus acquisition after an April NDRC order — the first forced unwind of a completed cross-border AI deal. A new exit template emerges.