Anthropic: How We Contain Claude ...

Anthropic: How We Contain Claude Across Products

AI
The Daily Diff by Premchand Chidipoti
S1 · E12
Aug 20, 2026
08:44

Episode notes

A rare, honest look at agent security from a lab shipping agents at scale. Jordan and Riley get into the central idea — "blast radius" (likelihood of failure x damage per failure): safeguards keep pushing likelihood down, but the worst-case damage only grows as agents gain capability and access, so the engineering job becomes bounding the damage, not preventing every failure. They cover why human-in-the-loop degrades (Anthropic's telemetry: users approved ~93% of Claude Code permission prompts -> approval fatigue), the shift to containment (sandboxes, VMs, egress controls), three risk types (user misuse, model misbehavior, external attackers), and the three containment shapes for claude.ai / Claude Code / Cowork matched to how much oversight each user can give. Best lessons: "the software you build yourself is the weakest" (their custom allowlist proxy failed while stock hypervisor/gVisor/seccomp held); the egress incidents where the model layer had nothing anomalous to catch; VM isolation locking EDR out too; tool/MCP output as a prompt-injection surface; and forward risks like persistent memory poisoning (CLAUDE.md, agent state dirs), multi-agent trust escalation, and agent identity. Source: How we contain Claude across products — Anthropic Engineering Blog, 2026, by Max McGuinness, Mikaela Grace, Jiri De Jonghe, Jake Eaton & Abel Ribbink — https://www.anthropic.com/engineering/how-we-contain-claude This is commentary/summary in the hosts' own words, not a reproduction of the article.

Keywords

Tech blog
Engineering blog
Software design
Software engineering