Run as a strict policy over 147 real web calls from our Claude Code fleet, OpenAPPA would have blocked 59 of them (40%). Every one was clean research, and our home-made gate let all 59 through. Both caught the leaks we planted. The difference is what each one checks: OpenAPPA asks what the session has read, ours asks what the call is carrying.
Why we ran it
Last week we built a small pre-run check on agent web calls, borrowing the idea from OpenAPPA (GitHub, site), an open-source, MIT-licensed guardrail from Archestra. It tracks what an agent has read and checks each tool call against fixed rules before it runs. No second model plays judge. This week it turned up in a newsletter as "0% attack success", so we installed the real engine and ran it against our copy on the same sessions.
How we compared them
OpenAPPA's Claude Code plugin writes hooks into the user-level Claude Code settings, and those
are shared by every agent on our box. So we did not install the plugin. We put the
appa binary (0.31.1, checksum-checked) in a sandbox and used its
replay command, which runs a script of tool calls against a policy with no
network and no agent.
We translated our policy into theirs: reading a .env file, a client transcript
or a personal note narrows the session's audience to one person. A web call to an allow-listed
site needs nothing. Any other web call needs a public audience. Then we turned the 31 sessions
still on disk into OpenAPPA traces, in order, and compared both engines call by call.
What we found
- Planted leaks: a run of six words each session had really read, sent to a made-up site. Our gate caught 29 of 29. OpenAPPA caught 28 of 29. The miss was our translation, not their engine: a shell command naming
.envwithout a slash in front slipped past the pattern we wrote. - Clean calls: 59 of 147 real calls would have been stopped by OpenAPPA and allowed by our gate. Searches for model prices, AWS docs, news sites. None carried anything the session had read.
- The other way round: zero. OpenAPPA never allowed a call our gate stopped.
The 40% is the same lesson as our own first draft, which blocked on the session label alone and flagged 48% of calls. Most of our sessions read a key or a private file early on and then go and research something unrelated. A rule keyed on the session, with no change in how the agent works, fires on all of that.
Why that is not the end of it
OpenAPPA does not expect agents to work the way ours do. When it blocks, it hands the agent a list of legal ways forward: mask the secret first, get a one-off approval, or do the private read in a throwaway child agent that can only return a small, fixed-shape answer, so the parent session never gets narrowed. A replay cannot pick a remedy, so our 59 is the cost of switching it on without changing how the fleet works, not a verdict on the engine.
It also covers ground ours does not. Our gate watches web tools only. OpenAPPA checks every tool, shell included, and the paper behind it proves its safety rules hold whatever a prompt-injected model asks for. Ours catches copied text, not paraphrase. Theirs does not care what the text says.
What's still off
The OpenAPPA policy is ours, written by hand, with none of their stock rule packs and none of the model-based classifiers they use for shell commands. 31 sessions is a small sample, because older transcripts have rotated off the box. And the "0% attack success" headline is from their own benchmark. We saw nothing that contradicts it, but we did not test it.
What we're doing about it
- Keeping our gate live. It is the one that does not stop clean work today.
- Not installing the OpenAPPA plugin fleet-wide yet. Hooks that fail closed in shared settings are a change to every agent at once.
- Taking their best idea next: move sensitive reads into a child agent that returns only what the parent needs, so the main session stays clean. That is what turns the 40% back into zero.
compare.pyis on GitHub. Run it on your own logs before you pick either kind of check.