Our nightly eval loop caught a bug, opened tickets for it, and we fixed the bug at the root. Then the loop kept re-picking those same tickets on later nights, because nothing ever retired them. Tonight's build closes that gap: a reconciler that removes a finding once its whole class is provably dead, with two guards so a still-live problem is never buried. On the first run it retired 6 stale flags.
What we did
Vera, our standing eval, grades what the fleet shipped each night and
opens a Gary to-do whenever an output scores below bar. Good behaviour for
real deliverables. The trouble starts when the fix is structural. Back on
1 August we stopped the panel from grading its own bookkeeping rows
(the per-juror votes and contract fingerprints it writes to the same
audit log) by excluding those action types in one place,
vera.rubric_gen.SELF_EVAL_ACTIONS. That class of flag can no
longer be raised. But the to-dos raised before the fix stayed open, so the
autonomous build slot kept selecting them. We almost re-fixed an
already-fixed bug tonight.
So we built vera.reconcile_flags. It scans open flag to-dos
and retires one only when both guards hold. Guard one is structural: the
flagged action is now in SELF_EVAL_ACTIONS, so the grader
cannot raise it again. Guard two is empirical: no flag of that same action
has been created on or after the fix landed. If a fresh one exists, the
class is not actually dead, so every flag of that action is left open and
the run says so. It is dry-run by default, idempotent, and refuses to
touch anything it cannot prove dead.
Why it was worth doing
A self-improving loop is only as good as the way it closes findings. If it
opens work but never retires resolved work, the backlog fills with dead
tickets and the loop spends its scarce nightly build slot re-doing settled
bugs. That is not a hypothetical: seven poll_juror and poll_contract flags
were sitting open days after their fix shipped, and two clean nightly runs
had already proven the class silent. The reconciler turned that proof into
action and cleared 6 of them in one pass (the seventh was the item we were
formally working, closed by hand). It is now wired into the same
standing-cron.sh right after the eval opens each night's
flags, so a class that goes quiet is cleaned the same night.
What's still off
This only retires the self-eval telemetry classes, the flags Vera raises about its own instrumentation. It deliberately never touches a flag about a real agent deliverable (a genuine bad answer from a working tool), which still needs a human or a build to resolve. It also keys off an explicit "fix landed" timestamp we set by hand when the exclusion shipped; a future version should read that from the change record rather than a constant. And it trusts the two-guard test rather than re-running the underlying repro, which is the right trade for telemetry but would not be for deliverables.
What's now in the stack
python -m vera.reconcile_flags— dry-run: shows which flags would retire and which classes are still live.python -m vera.reconcile_flags --apply— retires the proven-dead flags, with a reason written to each to-do.build_plan()— the pure, I/O-free decision function, covered byrepro_reconcile_flags.py(dead retires, live kept, deliverables untouched, idempotent).- Wired into
vera/standing-cron.shto run nightly after the eval, non-fatally.