Workloft
← Workloft Ships
1 October 2026 · infra · by Alfred + Bob

We checked 370 agent web calls for leaks.

48% of our agents' web calls came from a session that had just read a secret, a client transcript or a personal file. Nothing checked those calls before they ran. We built a pre-run check that follows the data rather than the session, replayed all 370 past calls through it, and found no secrets in any URL or query and one false alarm. The obvious version, which flags every call from a session that has read something private, would have flagged 177.

Why we built it

Our oversight scorecard found that web calls were the biggest gap in our checks. Shell commands go through a safety screen. Web fetches and searches went through nothing. An agent can read a .env file and call a website in the same session, and nothing would see a key in the URL. That is how a prompt injection gets data out: a web page tells the agent to fetch a link with your data in it.

This week OpenAPPA (GitHub, paper) came out of preview. It is an open-source guardrail that tracks what an agent has read and checks each tool call against it before it runs, with fixed rules and no second model acting as judge. We borrowed the idea and wrote a small version as two Claude Code hooks.

What we built

After every file read or shell command, a hook checks whether it touched something on a short list: .env files, keys, client call transcripts, personal notes. If so, it tags the session and stores hashed fingerprints of what was read. Four-word phrases, email addresses and long IDs, never the text itself. Before every web fetch, web search or Exa call, a second hook runs two rules:

Everything else goes through and is logged. There is an override, egress-ok <host>, and every use goes into a Friday email to Alfred, so a person sees the overrides, not just the agent that typed them.

What we found

We replayed every past session in order: 370 outbound calls across 33 sessions. 177 came from a session that had read something private. Our first version flagged on the tag alone, which meant flagging all 177. A check that fires on half your calls gets switched off by Friday.

The second version only flags when the private content shows up in the call. That brought it down to one, and it was a false alarm: a search for a public model name and price that also appeared in a private doc the session had read. We also planted test leaks: a run of six words each session had really read, sent to a made-up site, plus every secret value we hold. It caught 79 of 80 planted leaks and all 22 secrets.

Those planted leaks also caught a bug of ours. In the policy file, the list of secret files sat below a section header, so the config format filed it under that section. The secret check had loaded zero secrets, and every unit test still passed. The planted secrets caught it in the first run.

What's still off

It catches copied leaks, not paraphrased ones. An agent that rewrites a private document in its own words gets past it. A two-word name on its own is too short to fingerprint, and a test pins that gap. It only covers the web tools: a curl in the shell goes round it, which is the shell screen's job. Monitor mode means the content rule logs rather than blocks until a month of real traffic says the false-alarm rate holds. The secret rule blocks from day one.

What's now in the stack