Dev

Claude Managed Agents web_fetch: a URL now has to come from the user or the web

The new rule does not rank sources by how far you trust them. An attacker's web page counts as a source and your own bash output does not, because the model drives bash and does not write the web. With the web tool fenced, the next way out on a default environment is curl.

Since October 7, the web_fetch tool in Claude Managed Agents fetches only a URL that has already appeared in the session. The release note gives the list and the reason:

text
the web_fetch tool now fetches only URLs that have already appeared in the
session, for example in the text of a user message, in a web_search result,
or in a page that web_fetch returned earlier. This reduces the risk of data
exfiltration.

The same note lists what does not count: a URL that appears only in Claude's own output, the agent's system instructions, an attached document, or the output of bash, read or an MCP tool. Fetching one returns a url_not_in_prior_context error. To hand the agent a URL on purpose, put it in the text of a user.message.

My read is that the rule matters less than what it leaves alone. On an environment created with the API defaults, the URL it refuses is still one shell command away.

What does the Claude Managed Agents web_fetch rule actually block?

It blocks the plainest exfiltration an injected agent has: write a secret into a URL and fetch it. An instruction planted in a file or a page can no longer get https://collector.example/?k=<your key> out through web_fetch, as long as no outside source ever showed that URL. The attacker can write a link on a page, but cannot know your key in advance to put it there. The note does not say whether the match is on the full URL or something looser, and the protection depends on it being the full URL.

Why does a web page count as prior context in Managed Agents and bash output not?

Not because of trust. A fetched page is some of the least trustworthy text in a session, and it counts. Your bash output is produced inside the sandbox by a tool you enabled, and it does not. On my reading, the line falls on whether the model can drive the channel. In Managed Agents, bash, read and MCP calls are made by the model inside the session, so one echo would turn any URL Claude composes into "tool output" and make the rule decorative. A search result or a fetched page comes back from outside, already written. The system instructions and attached documents are excluded too, though the model cannot write them, so the rule errs toward refusing.

The Messages API version of the check fits the same reading. There, results from your client-side tools count as prior context "even when they echo text that Claude produced," because your code ran the tool and admitted the result. Server-side tools such as code execution and the MCP connector do not count. Managed Agents runs the sandbox for you, so its bash lands on the server side. The trust ordering is odd, since a web page outranks your own tools, but the decision is the kind I argued for in the connector essay: made by the channel a span came through, which the harness knows, not by reading the span.

What changed in Claude Managed Agents limited networking the same day?

Two entries landed with it. A cloud environment with limited networking now applies its allowed_hosts to web_search and web_fetch too, and with no hosts listed neither tool returns a page or a result. Creating a session, or updating one, now fails with a 400 when an enabled web tool's allowed_domains has an entry outside allowed_hosts. The lists match differently: a tool entry covers subdomains, an allowed_hosts entry is one exact host unless it starts with *., so docs.example.com is not within ["example.com"].

The side effect is the part to plan for. The environment docs say a host you add to allowed_hosts for the web tools "is also open to the sandbox," and that access is granted "per host, not per operation," uploads included. Every site you open for your agent to read under limited is also a site it can curl to.

What still leaves a Claude Managed Agents session after this change?

bash, if you let it. An environment created through the API without a networking field gets unrestricted, and the agent toolset's default permission policy is always_allow. On those defaults, a URL refused only by the prior-context check can still be requested from the shell, with nothing pausing it beyond a general safety blocklist. The new rule closes one door in an open building. That shape is not new: in OpenAI's misalignment reports, browsers refused URLs while uploads from the shell went through.

Two gaps remain even with the network locked. A page the agent fetched can carry many prepared links, and which one it follows is itself a signal; if the destination is allowed and the attacker can see requests to it, each fetched page can serve the next set of links, slow but with no ceiling. That is my reading, and nothing in the docs addresses it. And user.message is trusted on its face: if your app relays end-user text into it, every URL in that text passes the prior-context check, with domain lists still applying. The auto permission policy reads user.message as your intent the same way. It is an automated judgment, not a review, and the same caution applies as with Claude Code's approval classifier.

The docs leave a few things open. The Messages API refuses a URL that appears to carry a credential not shown in the system instructions or the user's text; the Managed Agents note does not say whether that check applies. Nor does it say where custom tool results fall, though they come back from your code, like client tools in the Messages API.

How should you configure a Claude Managed Agents environment on day one?

Set networking to limited explicitly, every time, and list only the hosts the job needs. If the agent only reads the web, disable bash so the shared host list opens nothing to a shell. If it needs one, put bash on always_ask; auto pauses sometimes, which beats never, but nobody looks before an allowed call runs. Pass the URLs you want fetched in a user.message, now the documented way in, and keep untrusted user text out of it unless you mean to grant it the same standing. The rule settles which URLs the model may not type. It says nothing about what else in the sandbox can.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles