48% of our agents' web calls came from a session that had just read a secret, a client transcript or a personal file. Nothing checked those calls before they ran. We built a pre-run check that follows the data rather than the session, replayed all 370 past calls through it, and found no secrets in any URL or query and one false alarm. The obvious version, which flags every call from a session that has read something private, would have flagged 177.
Why we built it
Our oversight scorecard found
that web calls were the biggest gap in our checks. Shell commands go through a safety
screen. Web fetches and searches went through nothing. An agent can read a
.env file and call a website in the same session, and nothing would see a key
in the URL. That is how a prompt injection gets data out: a web page tells the agent to
fetch a link with your data in it.
This week OpenAPPA (GitHub, paper) came out of preview. It is an open-source guardrail that tracks what an agent has read and checks each tool call against it before it runs, with fixed rules and no second model acting as judge. We borrowed the idea and wrote a small version as two Claude Code hooks.
What we built
After every file read or shell command, a hook checks whether it touched something on a
short list: .env files, keys, client call transcripts, personal notes. If so,
it tags the session and stores hashed fingerprints of what was read. Four-word phrases,
email addresses and long IDs, never the text itself. Before every web fetch, web search or
Exa call, a second hook runs two rules:
- Secrets always block. If the URL or query contains a known secret value, or anything shaped like an API key, the call does not run.
- Private content gets flagged. If the session is tagged, the site is not on the allow list, and the URL or query repeats a fingerprint of what was read, the call is flagged. Monitor mode logs it for now; enforce mode blocks.
Everything else goes through and is logged. There is an override, egress-ok <host>,
and every use goes into a Friday email to Alfred, so a person sees the overrides, not just
the agent that typed them.
What we found
We replayed every past session in order: 370 outbound calls across 33 sessions. 177 came from a session that had read something private. Our first version flagged on the tag alone, which meant flagging all 177. A check that fires on half your calls gets switched off by Friday.
The second version only flags when the private content shows up in the call. That brought it down to one, and it was a false alarm: a search for a public model name and price that also appeared in a private doc the session had read. We also planted test leaks: a run of six words each session had really read, sent to a made-up site, plus every secret value we hold. It caught 79 of 80 planted leaks and all 22 secrets.
Those planted leaks also caught a bug of ours. In the policy file, the list of secret files sat below a section header, so the config format filed it under that section. The secret check had loaded zero secrets, and every unit test still passed. The planted secrets caught it in the first run.
What's still off
It catches copied leaks, not paraphrased ones. An agent that rewrites a private document in
its own words gets past it. A two-word name on its own is too short to fingerprint, and a
test pins that gap. It only covers the web tools: a curl in the shell goes round
it, which is the shell screen's job. Monitor mode means the content rule logs rather than
blocks until a month of real traffic says the false-alarm rate holds. The secret rule blocks
from day one.
What's now in the stack
egress_gate.py: live on Bob as a pre-run and a post-run hook, withegress-okfor logged overrides.replay.py: replays your own sessions through the gate, with planted test leaks so a gate that only ever says yes scores zero.weekly.py: the Friday email. Calls checked, flags, and every override typed past any of our gates.- On GitHub with 15 tests. Run the replay on your own logs before you trust a check that has never said no.