The agent leash is now open source
NVIDIA released a free, Apache-licensed sandbox that inspects everything an AI agent tries to send out, and a watchdog chip that can kill the agent even if the machine is owned. The same day, Anthropic made running those agents about 30% cheaper.

For two years the AI agent conversation has been stuck in a loop. Someone demos an agent that books a flight, files a ticket, refactors a codebase. Someone else asks the obvious question: what stops it? And the answer, embarrassingly often, has been a paragraph of polite English in a system prompt. "Do not access anything outside the approved list." That is not a control. That is a note on the fridge.
On 28 September that changed in a way I did not expect. NVIDIA open-sourced the leash.
What actually shipped
The announcement is called the Open Agent Safety Platform, and NVIDIA calls it an "open reference design," which is a modest label for something with this much weight behind it. It comes in two halves.
The software half is OpenShell, released under the Apache 2.0 licence on GitHub. Per SecurityWeek's write-up it has three moving parts: a gateway that manages sandbox lifecycles and policy, a sandbox with kernel-level controls over filesystem and process activity, and a supervisor that sits next to each sandbox and checks every outbound request against policy. You can, for instance, let an agent read through an API while blocking every write. Crucially, the controls stay on even when the agent starts executing code it wrote itself, which is exactly the moment most homemade guardrails quietly stop applying.
The detail I keep coming back to is the credentials. In OpenShell, the agent only ever sees a placeholder API key. The real one gets swapped in outside the sandbox, and only for endpoints on the approved list. Think about what that does to the entire prompt-injection genre. If a poisoned web page talks your agent into exfiltrating its own token, it exfiltrates a fake. The secret was never in the room.
There is also a "policy advisor" that lets the agent propose a permission change but never approve one, and a formal policy prover that checks granted permissions stay inside the operator's limits. Propose, not approve. Someone at NVIDIA has clearly watched an agent argue itself into more access.
The part that runs on a different chip
The hardware half is called Sentry, and it runs on BlueField-4 DPUs, physically separate from the host the agent lives on. NVIDIA's claim is that if an agent tries to cross its boundary, Sentry quarantines and stops it "in milliseconds" even if the host itself has been compromised.
Here is the analogy that made it click for me. Every software sandbox is a security guard who works for the building. If an attacker takes the building, they take the guard. Sentry is a guard standing in the street with his own radio, watching the doors from outside. He does not care what you have done to the lobby.
The check has moved from the agent's instructions to the path out. That is the whole story, and it is the only version of agent safety that survives contact with a clever attacker.
The catch is real: Sentry needs NVIDIA's Vera Rubin data-centre hardware, with existing Vera systems getting it by software update. Almost nobody reading this will own that. The software half is free and the hardware half is not for you, and both of those facts matter.
What tells you this is not a one-vendor experiment is the roster. NVIDIA says more than 100 organisations are involved, and names Anthropic, Microsoft, Salesforce, SAP, Red Hat, IBM, Palantir, Scale AI, Figure, SpaceXAI and Palo Alto Networks; SecurityWeek adds CrowdStrike and Cisco. Model labs, cloud vendors, security vendors and robotics companies do not agree on much. They agree on this.
The same day, agents got cheaper to run
Anthropic released Claude Sonnet 5.5 on 28 September at exactly the old Sonnet price: $2 per million input tokens, $10 output, $0.20 cache reads. The headline claim is up to 30% lower total cost per task, and the mechanism is the interesting bit. The rate did not drop. The model just does less flailing: fewer tokens, and fewer tool calls to finish the same job.
Anthropic's own benchmark numbers put Sonnet 5.5 at 1844 on GDPval-AA against Opus 5.5's 1846, up from Sonnet 5's 1449. Customers quoted in the coverage report the same shape: Lovable saw one-third fewer tool calls and roughly half as many shell executions, Base44 got builds done in 3.6 iterations versus 7.7 on Opus 5, Box measured 2.4x faster with 12% fewer total tokens.
Fair warning, and I want to be blunt about it: every one of those figures is vendor-reported. The Decoder notes plainly that independent testing has not confirmed the speed and cost claims. Treat them as a hypothesis with a price tag attached, not a result.
Why fewer tool calls is the number to watch
Cost is the boring reason to care. Reliability is the real one. Every tool call an agent makes is another chance to fail, loop, hallucinate a parameter or slam into a guardrail. An agent that finishes in three calls instead of seven is not just cheaper, it is a smaller attack surface and a shorter list of things that can go wrong at 2am.
Put the two announcements side by side and you get the actual signal from this week. The boundary around an AI agent stopped being a design philosophy and became a repo you can clone, a chip you can buy, and a line item that just got 30% smaller. Anthropic also shipped its first Sonnet with dedicated cybersecurity safeguards, visibly rerouting high-risk cyber requests to the older model rather than silently refusing them.
The question everyone was asking a week ago was "can we stop the agent?" The market's answer this week is yes, and here is the part number.
If you run agents in production and you are still relying on instructions rather than an enforced path out, the excuse just got a lot thinner. The tooling is free, the licence is permissive, and the people who build the models are in the room.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- NVIDIA Newsroom - Open Agent Safety Platform - https://nvidianews.nvidia.com/news/open-agent-safety-platform
- NVIDIA Technical Blog - a reference for continuous in-silicon agent monitoring - https://developer.nvidia.com/blog/nvidia-open-agent-safety-platform-a-reference-for-continuous-in-silicon-agent-monitoring/
- GitHub - NVIDIA/openshell - https://github.com/NVIDIA/openshell
- SecurityWeek - NVIDIA unveils AI agent safety platform with hardware-based watchdog - https://www.securityweek.com/nvidia-unveils-ai-agent-safety-platform-with-hardware-based-watchdog/
- VentureBeat - Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per task - https://venturebeat.com/technology/anthropic-launches-claude-sonnet-5-5-with-30-cost-reduction-per-task-due-to-faster-speeds-and-fewer-tool-calls
- The Decoder - Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task - https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/
Quick answers
What is NVIDIA's Open Agent Safety Platform?
An open reference design announced on 28 September 2026 with two parts. OpenShell is an Apache 2.0 licensed runtime on GitHub with a gateway, a sandbox using kernel-level filesystem and process controls, and a supervisor that checks every outbound request against policy. Sentry is a watchdog running on BlueField-4 DPUs, separate from the host, that NVIDIA says quarantines and stops a boundary-crossing agent in milliseconds.
What does the placeholder credential trick actually do?
In OpenShell the agent only ever sees a fake API key. The real credential is substituted outside the sandbox, and only for approved endpoints. If a prompt injection convinces the agent to leak its token, what leaks is worthless, because the secret was never inside the agent's environment.
Do I need NVIDIA hardware to use any of this?
Only for the Sentry half, which requires NVIDIA Vera Rubin hardware, with existing Vera systems getting it via software update. OpenShell, the software half, is open source under Apache 2.0 and is the part most teams can actually adopt.
Is Claude Sonnet 5.5 really 30% cheaper?
Anthropic claims up to 30% lower total cost per task and more than 30% faster output, achieved through fewer tokens and fewer tool calls rather than a lower price per token. The list price is unchanged at $2 input and $10 output per million tokens. These are vendor-reported figures and independent testing has not confirmed them.