Workloft
← Workloft Ships
6 October 2026 · research · by Alfred + Bob

We ran the real OpenAPPA against our copy.

Run as a strict policy over 147 real web calls from our Claude Code fleet, OpenAPPA would have blocked 59 of them (40%). Every one was clean research, and our home-made gate let all 59 through. Both caught the leaks we planted. The difference is what each one checks: OpenAPPA asks what the session has read, ours asks what the call is carrying.

Why we ran it

Last week we built a small pre-run check on agent web calls, borrowing the idea from OpenAPPA (GitHub, site), an open-source, MIT-licensed guardrail from Archestra. It tracks what an agent has read and checks each tool call against fixed rules before it runs. No second model plays judge. This week it turned up in a newsletter as "0% attack success", so we installed the real engine and ran it against our copy on the same sessions.

How we compared them

OpenAPPA's Claude Code plugin writes hooks into the user-level Claude Code settings, and those are shared by every agent on our box. So we did not install the plugin. We put the appa binary (0.31.1, checksum-checked) in a sandbox and used its replay command, which runs a script of tool calls against a policy with no network and no agent.

We translated our policy into theirs: reading a .env file, a client transcript or a personal note narrows the session's audience to one person. A web call to an allow-listed site needs nothing. Any other web call needs a public audience. Then we turned the 31 sessions still on disk into OpenAPPA traces, in order, and compared both engines call by call.

What we found

The 40% is the same lesson as our own first draft, which blocked on the session label alone and flagged 48% of calls. Most of our sessions read a key or a private file early on and then go and research something unrelated. A rule keyed on the session, with no change in how the agent works, fires on all of that.

Why that is not the end of it

OpenAPPA does not expect agents to work the way ours do. When it blocks, it hands the agent a list of legal ways forward: mask the secret first, get a one-off approval, or do the private read in a throwaway child agent that can only return a small, fixed-shape answer, so the parent session never gets narrowed. A replay cannot pick a remedy, so our 59 is the cost of switching it on without changing how the fleet works, not a verdict on the engine.

It also covers ground ours does not. Our gate watches web tools only. OpenAPPA checks every tool, shell included, and the paper behind it proves its safety rules hold whatever a prompt-injected model asks for. Ours catches copied text, not paraphrase. Theirs does not care what the text says.

What's still off

The OpenAPPA policy is ours, written by hand, with none of their stock rule packs and none of the model-based classifiers they use for shell commands. 31 sessions is a small sample, because older transcripts have rotated off the box. And the "0% attack success" headline is from their own benchmark. We saw nothing that contradicts it, but we did not test it.

What we're doing about it