fix(scanner): enforce the verdict schema where the rows are written - #157
Merged
Merged
Conversation
ADR-016 shipped the dismissal corpus on 2026-09-21. Its first production run, scan #109 on 2026-09-22, produced verdicts.json with 7 scanner rows and 6 curation rows — the plumbing worked — and every curation row was invalid. All six had an empty `rule`; all six carried free prose in `verdict` ("by-design not-reachable", "FP - parameterized", "real - lockfile refresh"). The unit tests passed the whole time, because they exercise the constructor and production does not: the curation half is hand-written JSON, so VerdictRecord.__post_init__ never runs. A closed vocabulary is only closed if something closes it. Four fixes: - `scanner/verdict_corpus.py` + `scanner/run_verdict_check.py`. `--check` reports every problem in every row in one pass and exits 2; `/daily` Phase 4 must see exit 0 before publishing. A checker that reveals one defect per run costs as many runs as there are defects. - `rule` is now required on every record — ADR-014 needed "sqlalchemy-execute-raw-query fired 157 times", not "some rules were dismissed". - `dependency-currency` joins the vocabulary; `confirmed-real` is documented as first-party only. Scan #109 filed 13 dependency CVEs as `confirmed-real` under a headline of zero real findings. - `corpus/verdicts/` is committed. `reports/` is gitignored — it holds raw dumps and unsent disclosure drafts — so the corpus lived on one machine and would not survive a clone. `--summary` counts it. `scanner/verdicts.py` hit 296 of the 300-line ceiling, so validation and corpus loading moved to the new module. Scan #109's six invalid rows are left alone: rewriting them would mean inventing rule identities that were never recorded, and a corpus with one fabricated entry is worth less than one with a known hole. ADR-017. 26 new tests; 416 pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #154 (ADR-016), driven by its first production run.
What the first run showed
Scan #109 (
overwirehq/claude-code-telegram, 2026-09-22) producedverdicts.jsonwith 7 scanner rows and 6 curation rows — the plumbing worked end to end, and the published post carried both new blocks.Every one of the 6 curation rows was invalid:
ruleemptyverdictfree prose ("by-design not-reachable","FP - parameterized")confirmed-realused for 13 dependency CVEs under a "0 real" headlineThe test suite was green throughout. It exercises the constructor; production does not — the curation half is hand-written JSON, so
VerdictRecord.__post_init__never runs. A closed vocabulary is only closed if something closes it.A fourth problem is structural rather than drift:
reports/is gitignored, so the corpus lived on one machine and would not survive a clone. ADR-014's measurement only worked because the archived reports happened to still be on disk.Fixes
scanner/verdict_corpus.py+scanner/run_verdict_check.py—--checkreports every problem in every row in one pass and exits 2./dailyPhase 4 must see exit 0 before publishing. A checker that reveals one defect per run costs as many runs as there are defects, so fields are checked independently rather than through the first-raising constructor.rulerequired on every record. ADR-014 neededsqlalchemy-execute-raw-query fired 157 times, not "some rules were dismissed".dependency-currencyadded;confirmed-realdocumented as first-party only.corpus/verdicts/committed — the durable half.reports/stays ignored (it holds raw dumps and unsent disclosure drafts); a negation rule there would need a fragilereports/*+!reports/*/ladder that risks exposingreports/disclosures/.--summarycounts the corpus.scanner/verdicts.pyhit 296 of the 300-line ceiling, so validation and corpus loading moved out.Not fixed, on purpose
Scan #109's six invalid rows are left as they are. Repairing them would mean inventing rule identities that were never recorded, and a corpus with one fabricated entry is worth less than one with a known hole. Scan #110 is the first valid entry.
Verification
ruffclean ·black77 files unchanged ·pytest416 passed (26 new)--checkrun against the real scan docs: seventy-ninth public scan (liaohch3/claude-tap) — 1 real finding withheld #109 file: 12 problems across 6 rows, in one pass🤖 Generated with Claude Code