Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/skills/daily/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ Update outcomes ONLY after upstream merge/close, never on open. Stop here if `--
Then read `reports/<slug>/coverage.json` — the authoritative record of what the scan examined, written every run from the scanners' own meta findings and derived **before** `--ignore-file` suppression. If `complete` is `false`, the finding count is not a clean bill of health: carry the rows verbatim into the post's **Scan coverage** section and state in one sentence what was not examined. This replaces transcribing the Phase 1 preflight by hand; the preflight still runs, because catching a broken Semgrep *before* burning a scan beats reporting it afterwards.

## Phase 4 — Curate
Group by rule family; auto-flag tests/sample/examples/demos/fixtures as candidate-FP. Inspect top 5 real candidates in-repo via `gh api .../contents/<path>`. Write per-finding verdicts, then append them to `reports/<slug>/verdicts.json` as `source: "curation"` rows — one per rule family with its count, beside the deterministic rows the scanner already wrote. `reason_code` must come from the closed vocabulary in `scanner/verdicts.py:CURATION_REASON_CODES`; reuse or extend the module, never invent a code inline. Evaluate the quality gate boolean.
Group by rule family; auto-flag tests/sample/examples/demos/fixtures as candidate-FP. Inspect top 5 real candidates in-repo via `gh api .../contents/<path>`. Write per-finding verdicts, then append them to `reports/<slug>/verdicts.json` as `source: "curation"` rows — one per rule family with its count, beside the deterministic rows the scanner already wrote. `reason_code` must come from the closed vocabulary in `scanner/verdicts.py:CURATION_REASON_CODES`; reuse or extend the module, never invent a code inline. `rule` is REQUIRED (name the rule family); `verdict` is one of false-positive/by-design/not-applicable/hardening/real or omitted, and your sentence goes in `detail`; `confirmed-real` is first-party only, upstream version drift is `dependency-currency`. Then validate — `.venv/Scripts/python.exe scanner/run_verdict_check.py --check reports/<slug>/verdicts.json` must exit 0 — and copy the validated file to `corpus/verdicts/<slug>.json`, which is committed (`reports/` is gitignored). Evaluate the quality gate boolean.

## Phase 5 — Publish (gated)
Always: write `docs/scans/<slug>.md` (including the **Scan coverage** block copied from `reports/<slug>/coverage.json`, never hand-written) + prepend a row to `docs/index.md` + a bullet to `docs/scan-log.md` (bump both scan counts). If gate TRUE and not strict-norm: file focused de-branded courtesy issue (+ PR for clean one-line fixes). If strict-norm: post-only or one-issue-per-critical. If gate FALSE: post-only. Open `docs:` PR on `elfrost/ai-patchlab`, merge, verify publication via `gh run list --workflow=pages-build-deployment --limit 1` = `success` AND the post returns 200 serving the new text. Do NOT gate on `pages/builds/latest` — on this Actions-published site it reports `errored` even when the deploy succeeded (2026-08-26).
Expand Down
9 changes: 8 additions & 1 deletion .claude/commands/daily.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,8 +119,15 @@ Then read `reports/<slug>/coverage.json`. It is the authoritative record of what
2. Inspect the top 5 real candidates in the actual repo via `gh api repos/<owner>/<name>/contents/<path>` — read the call site, confirm the threat path.
3. Write per-finding verdicts (real / by-design / FP, with the *why*), then **append them to `reports/<slug>/verdicts.json`** as `source: "curation"` rows — one row per rule family with its count, not one per finding. The scanner has already written its own deterministic rows there (what `--ignore-file` and `--min-severity` removed); add yours beside them.
`reason_code` comes from the closed vocabulary in `scanner/verdicts.py:CURATION_REASON_CODES` — `sql-identifier-fp`, `test-or-fixture-path`, `sample-or-demo`, `vendored-code`, `not-reachable`, `mitigated-in-app`, `by-design`, `product-surface`, `domain-noun-collision`, `placeholder-secret`, `active-harm-fp`, `credited-defense`, `confirmed-real`. Reuse a code or add one to the module; never invent one inline, because a long tail of one-off codes counts to one and the corpus stops being countable.
`rule` is REQUIRED on every row — name the rule family (`github-actions-mutable-action-tag`, `sqlalchemy-execute-raw-query`, the CVE id). A row without it can be counted but never acted on, which defeats the file. `verdict` is one of `false-positive` / `by-design` / `not-applicable` / `hardening` / `real`, or omitted — **your sentence goes in `detail`, never in `verdict`**. `confirmed-real` is first-party only; an upstream dependency merely behind a fixed version is `dependency-currency`.
4. **Validate the file before publishing.** It is hand-written JSON, so nothing else enforces the schema:
```bash
.venv/Scripts/python.exe scanner/run_verdict_check.py --check reports/<slug>/verdicts.json
```
Exit 0 or fix what it lists and re-run. On 2026-09-22 the first run wrote six rows with an empty `rule` and free prose in `verdict`, and nothing caught it — a closed vocabulary is only closed if something closes it.
5. Copy the validated file to `corpus/verdicts/<slug>.json` and commit it with the post. `reports/` is gitignored, so a corpus left there lives on one machine and does not survive a clone; `corpus/` is the durable half.
This is the file ADR-014 had to reconstruct by hand from 87 archived reports. Writing it as you go is what turns "13th appearance of this FP" into a number that justifies mechanizing the rule.
4. **Evaluate the quality gate:** is there ≥1 real, exploitability-shaped, high-confidence item? Record the boolean — it decides Phase 5 filing.
6. **Evaluate the quality gate:** is there ≥1 real, exploitability-shaped, high-confidence item? Record the boolean — it decides Phase 5 filing.

## Phase 5 — Publish (gated)
1. **Always:** write `docs/scans/<slug>.md` from `docs/templates/scan-post.md` — including the **Scan coverage** block, copied from `reports/<slug>/coverage.json` and never hand-written; prepend a new row to the Scans table in `docs/index.md` **and** a new bullet to `docs/scan-log.md` (the full prose archive), and bump the scan counts in both headers. Three files, every time — on 2026-09-06 the log was found eight entries behind the index because this step only named `index.md`.
Expand Down
9 changes: 7 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,9 @@ AI PatchLab is an AI-assisted security remediation toolkit. The MVP focuses on a
- `scanner/models.py` - Normalized `Finding` dataclass + severity/confidence enums + `FINDING_FIELDS`
- `scanner/recommendations.py` - Deterministic keyword-based recommendation enrichment
- `scanner/coverage.py` - Per-scanner coverage derived from meta findings (`ToolCoverage`, `build_coverage`, `EXPECTED_TOOLS`, `is_complete`); feeds `reports/coverage.json` and the report's Scan Coverage block
- `scanner/verdict_corpus.py` - Corpus validation and aggregation (`validate_payload` reports every problem in one pass, `load_corpus` reads `corpus/verdicts/`)
- `scanner/run_verdict_check.py` - CLI: `--check <verdicts.json>` (exit 2 on any problem, required by `/daily` Phase 4) and `--summary` (count the whole corpus)
- `corpus/verdicts/` - Committed dismissal corpus, one file per scan; the durable half, since `reports/` is gitignored
- `scanner/verdicts.py` - Dismissal records (`VerdictRecord`, `summarize_removed`, `count_by_reason`, `load_records`) with closed `SCANNER_REASON_CODES` / `CURATION_REASON_CODES` vocabularies; feeds `reports/verdicts.json` and the report's Dismissed section
- `scanner/report_markdown.py` - Markdown rendering, split out of `report.py` to stay under the 300-line ceiling (`write_markdown_report` is still re-exported from `scanner.report`)
- `scanner/confidence.py` - Centralized `Finding.confidence` rules (one function per scanner + `confidence_for_meta_finding` for shared `not-installed` / `scan-error` / etc.)
Expand All @@ -74,7 +77,7 @@ AI PatchLab is an AI-assisted security remediation toolkit. The MVP focuses on a
- `src/main.py` - Legacy entry point (`python -m src.main`) - currently a loguru-wired async stub with TODOs
- `.github/workflows/ci.yml` - CI: ruff + black + pytest on Python 3.11 and 3.13
- `reports/disclosures/` - Drafted private disclosure emails awaiting a manual send (gitignored)
- `tests/` - pytest tests (one module per scanner: `test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`, `test_coverage.py`, `test_verdicts.py`)
- `tests/` - pytest tests (one module per scanner: `test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`, `test_coverage.py`, `test_verdicts.py`, `test_verdict_corpus.py`)
- `tests/conftest.py` - Shared fixtures (`mock_db`, `mock_http_client`, `mock_discord`, `test_config`, session `event_loop`)
- `examples/` - Reference patterns to read before implementing
- `PRPs/` - Active Product Requirements Prompts
Expand Down Expand Up @@ -181,6 +184,8 @@ export AI_PATCHLAB_AI_REVIEW_COMMAND=/path/to/ai-review-wrapper
- Do not call subprocesses directly from `scanner/scanners/*` - go through the runner module
- Every suppression step in `run_scan` records what it removed via `summarize_removed(before, after, reason_code)` - a narrower report must never shrink its own numbers silently (ADR-016)
- `reason_code` values are a closed vocabulary in `scanner/verdicts.py`; add a code to the module rather than inventing one at a call site, or the corpus stops being countable
- Curation rows are hand-written JSON, so the dataclass never runs on them - `/daily` Phase 4 MUST get exit 0 from `scanner/run_verdict_check.py --check` before publishing (ADR-017). `rule` is required on every row; `verdict` takes one of five values and the sentence goes in `detail`
- `confirmed-real` is first-party only; an upstream dependency merely behind a fixed version is `dependency-currency`
- Coverage is derived from the raw `collect_findings` output **before** `apply_ignore` (`scanner/run_scan.py`) - `--ignore-file` does not exempt meta findings, so deriving it later would let a path pattern hide the fact that a tool never ran
- Adding a scanner to `SCANNERS` requires adding its `Finding.tool` value to `scanner/coverage.py:EXPECTED_TOOLS`; `tests/test_coverage.py::TestRegistryDrift` fails until you do
- `Finding.confidence` values come from `scanner/confidence.py` - never inline `confidence="high"` / `"medium"` / `"low"` in a scanner adapter; add or reuse a rule function instead
Expand Down Expand Up @@ -235,7 +240,7 @@ export AI_PATCHLAB_AI_REVIEW_COMMAND=/path/to/ai-review-wrapper
- Log architectural decisions in `DECISIONS.md`
- Check existing ADRs before making structural changes
- Record date, decision, context, and consequences
- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact, ADR-016 dismissals recorded as counted rule families
- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact, ADR-016 dismissals recorded as counted rule families, ADR-017 a closed vocabulary needs a closer

## Known Gotchas
- **Do not tag scan posts by keyword inference.** A classifier over post bodies was built and rejected 2026-09-18: validated against 9 posts of known ground truth it gave klavis **6** finding classes where the real finding was dependency CVEs, and got 3 of 9 project families wrong (OpenBiliClaw as "developer tooling", tracecat as "MCP server"). Post bodies discuss false positives and credited defences at length, so matching them tags a clean scan with the class it *dismissed*. On a site whose whole argument is that pattern-matching produces plausible-but-wrong results, publishing plausible-but-wrong tags is self-refuting. Grouping pages are generated from the **curated index table** instead — hand-maintained, verified data
Expand Down
9 changes: 7 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,9 @@ This project can optionally include a parallel Codex/OpenAI runtime via `AGENTS.
- `scanner/models.py` — Normalized `Finding` dataclass + severity/confidence enums + `FINDING_FIELDS`; `Finding.is_meta` flags scanner-infrastructure findings that `--min-severity` must never drop
- `scanner/recommendations.py` — Deterministic keyword-based recommendation enrichment
- `scanner/coverage.py` — Per-scanner coverage derived from meta findings (`ToolCoverage`, `build_coverage`, `EXPECTED_TOOLS`, `is_complete`); feeds `reports/coverage.json` and the report's Scan Coverage block
- `scanner/verdict_corpus.py` — Corpus validation and aggregation (`validate_payload` reports every problem in one pass, `load_corpus` reads `corpus/verdicts/`)
- `scanner/run_verdict_check.py` — CLI: `--check <verdicts.json>` (exit 2 on any problem, required by `/daily` Phase 4) and `--summary` (count the whole corpus)
- `corpus/verdicts/` — Committed dismissal corpus, one file per scan; the durable half, since `reports/` is gitignored
- `scanner/verdicts.py` — Dismissal records (`VerdictRecord`, `summarize_removed`, `count_by_reason`, `load_records`) with closed `SCANNER_REASON_CODES` / `CURATION_REASON_CODES` vocabularies; feeds `reports/verdicts.json` and the report's Dismissed section
- `scanner/report_markdown.py` — Markdown rendering, split out of `report.py` to stay under the 300-line ceiling (`write_markdown_report` is still re-exported from `scanner.report`)
- `scanner/confidence.py` — Centralized `Finding.confidence` rules (one function per scanner + `confidence_for_meta_finding` for shared `not-installed` / `scan-error` / etc.)
Expand All @@ -70,7 +73,7 @@ This project can optionally include a parallel Codex/OpenAI runtime via `AGENTS.
- `src/main.py` — Legacy point d'entrée (`python -m src.main`) — currently a loguru-wired async stub with TODOs
- `.github/workflows/ci.yml` — CI: ruff + black + pytest on Python 3.11 and 3.13 (the 3.13 leg catches stdlib removals such as PEP 594 dropping `cgi`)
- `reports/disclosures/` — Drafted private disclosure emails awaiting a manual send (gitignored with the rest of `reports/`)
- `tests/` — Tests pytest (`test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`, `test_coverage.py`, `test_verdicts.py`)
- `tests/` — Tests pytest (`test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`, `test_coverage.py`, `test_verdicts.py`, `test_verdict_corpus.py`)
- `tests/conftest.py` — Fixtures partagées (`mock_db`, `mock_http_client`, `mock_discord`, `test_config`, session `event_loop`)
- `examples/` — Code de référence — LIRE AVANT D'IMPLÉMENTER (api_client, config, discord_alert, mysql, playwright_scraper, scheduler, service)
- `PRPs/` — Product Requirements Prompts (actifs)
Expand Down Expand Up @@ -121,6 +124,8 @@ This project can optionally include a parallel Codex/OpenAI runtime via `AGENTS.
- Any finding built with `confidence_for_meta_finding(...)` must also set `is_meta=True` so `--min-severity` cannot drop it
- Every suppression step in `run_scan` records what it removed via `summarize_removed(before, after, reason_code)` — a narrower report must never shrink its own numbers silently (ADR-016)
- `reason_code` values are a closed vocabulary in `scanner/verdicts.py`; add a code to the module rather than inventing one at a call site, or the corpus stops being countable
- Curation rows are hand-written JSON, so the dataclass never runs on them — `/daily` Phase 4 MUST get exit 0 from `scanner/run_verdict_check.py --check` before publishing (ADR-017). `rule` is required on every row; `verdict` takes one of five values and the sentence goes in `detail`
- `confirmed-real` is first-party only; an upstream dependency merely behind a fixed version is `dependency-currency`
- Coverage is derived from the raw `collect_findings` output **before** `apply_ignore` (`scanner/run_scan.py`) — `--ignore-file` does not exempt meta findings, so deriving it later would let a path pattern hide the fact that a tool never ran
- Adding a scanner to `SCANNERS` requires adding its `Finding.tool` value to `scanner/coverage.py:EXPECTED_TOOLS`; `tests/test_coverage.py::TestRegistryDrift` fails until you do
- `Finding.confidence` values come from `scanner/confidence.py` — never inline `confidence="high"` / `"medium"` / `"low"` in a scanner adapter; add or reuse a rule function instead
Expand Down Expand Up @@ -286,7 +291,7 @@ $env:AI_PATCHLAB_AI_REVIEW_COMMAND = "C:\tools\ai-review-wrapper.cmd"
- Before making a structural decision, check DECISIONS.md for precedent
- Use the architect agent (`/architect` or Task tool) for complex decisions
- Format: ADR (Architecture Decision Record) — date, decision, context, consequences
- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact, ADR-016 dismissals recorded as counted rule families
- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact, ADR-016 dismissals recorded as counted rule families, ADR-017 a closed vocabulary needs a closer

## Known Gotchas
- **Do not tag scan posts by keyword inference.** A classifier over post bodies was built and rejected 2026-09-18: validated against 9 posts of known ground truth it gave klavis **6** finding classes where the real finding was dependency CVEs, and got 3 of 9 project families wrong (OpenBiliClaw as "developer tooling", tracecat as "MCP server"). Post bodies discuss false positives and credited defences at length, so matching them tags a clean scan with the class it *dismissed*. On a site whose whole argument is that pattern-matching produces plausible-but-wrong results, publishing plausible-but-wrong tags is self-refuting. Grouping pages are generated from the **curated index table** instead — hand-maintained, verified data
Expand Down
Loading
Loading