Skip to content

feat: pattern-specific definitions and severity vectors for all 200 records - #10

Merged
requie merged 1 commit into
mainfrom
feat/pattern-definitions-and-severity
Sep 30, 2026
Merged

requie merged 1 commit into
mainfrom
feat/pattern-definitions-and-severity

Conversation

@requie

@requie requie commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Every record's definition was one template sentence ("Tests whether X can cross an agent trust boundary..."), and every record carried the same six-axis severity vector regardless of its score, so the vector explained nothing and the 175 non-reference scores were flat per-domain defaults. Each of the 200 patterns now has its own definition stating the mechanism, the trust boundary it crosses, and the unsafe result, and its own vector on the six CAAP axes (impact, exploitability, privilege, autonomy, persistence, propagation), each scored 1 to 5.

This is CAAP's own six-axis vector as defined in v1.0, not CVSS. No CVSS strings are introduced.

Pattern or implementation impact

Standards review: this changes the definition and severity of every record. Stable IDs, titles, families, maturity, mappings, and relationships are unchanged. The 25 v1.0 reference baseline scores are unchanged and marked score_source: caap-v1.0-baseline.

  • scripts/generate_catalog.py: new PATTERN_DETAILS table with a definition and vector per pattern, keyed by name; the generator refuses to build if any pattern is missing or any entry is orphaned. severity_from_vector derives the baseline score as 10 * (2*impact + 2*exploitability + privilege + autonomy + persistence + propagation) / 40, rounded to one decimal. severity_rating adds a low band below 4.0. Reference vectors must land within 0.5 of the preserved v1.0 score or generation fails; all 25 land within 0.2. Pattern pages gain a Severity line.
  • scripts/validate_repository.py: checks definitions are unique and non-template, vectors complete and in range, ratings match scores, non-reference scores equal their vector, and reference scores are the v1.0 baseline within tolerance.
  • tests/test_taxonomy.py: pins the 25 v1.0 scores as a regression guard; checks vector derivation, range, and distinctness; checks definitions are unique, specific, and at least 80 characters.
  • docs/STANDARD.md: new Severity section defining each axis, the formula, the rating bands (which match the 2.0.0-draft.1 standard document: critical 9.0+, high 7.0 to 8.9, medium 4.0 to 6.9, low below 4.0), the reference-score rule, and the caveat that baseline scores are illustrative estimates.
  • CHANGELOG.md: Changed entry.
  • Regenerated: data/taxonomy/caap-200.json and .yaml, docs/TAXONOMY.md, 200 pattern pages, 200 case files (severity is copied into each case), and the packaged copies. The diff is 433 files for that reason; only the generator, validator, tests, standard, and changelog are hand-edited.

Distribution after the change: 13 critical, 166 high, 21 medium, 0 low; 131 distinct vectors; scores 6.0 to 9.4. Before: 9 critical, 190 high, 1 medium; 1 vector.

Safety impact

Not applicable. No fixture, capability, sink, network boundary, persistence boundary, or public procedure changes. Definitions describe mechanisms at the level of the existing pattern names and carry no payloads or procedures.

  • Authorized targets only
  • Synthetic data and identities only
  • Mock tools, sinks, and actuators only
  • No destructive payload, real exfiltration, persistence, or approval bypass

Validation

  • Generated files are current (regenerated twice; no drift)
  • Repository validation passes, including full JSON Schema validation, and fails as intended when a score is tampered with
  • Unit tests pass (46 tests)
  • Secure mock passes (25 pass); the new scores flow into the severity-weighted report
  • Vulnerable synthetic behavior fails when an executable case changes (25 fail; oracles and traces unchanged)
  • ruff check src tests scripts examples is clean
  • Commit includes DCO sign-off

Review focus

The definitions and axis judgments are the substance of this PR and were drafted for domain-editor review, not just CI. Worth a second opinion: the four new critical ratings outside the reference set (Dependency Confusion, Build-Script Injection, CI Command Injection, Dependency Installation Hijack), the embodied-agent domain where physical harm is treated as high persistence, and the human-trust domain, which scores lowest on average because those mechanisms need a human to act. Correcting any vector is an edit to the generator's table plus regeneration.

…ecords

Every record's definition was one template sentence, and every record
carried the same six-axis severity vector regardless of its score, so
the vector explained nothing and 175 scores were flat per-domain
defaults. Each of the 200 patterns now has its own definition stating
the mechanism, the trust boundary it crosses, and the unsafe result,
and its own vector on the six v1.0 axes, each scored 1 to 5.

The baseline score is derived from the vector by a documented method
in docs/STANDARD.md, weighting impact and exploitability double. The
25 v1.0 reference scores are unchanged and marked caap-v1.0-baseline;
their vectors land within 0.2 of the preserved score and the generator
refuses anything beyond 0.5. All other scores are vector-derived.
Ratings gain a low band below 4.0. Pattern pages show the severity.

Repository validation now checks that definitions are unique and not
the template, vectors are complete and in range, ratings match scores,
non-reference scores equal their vector, and reference scores are the
v1.0 baseline within tolerance. Tests pin the 25 v1.0 scores as a
regression guard. Distribution after the change: 13 critical, 166
high, 21 medium, 131 distinct vectors, scores 6.0 to 9.4.

Signed-off-by: requie <tarique.smith@gmail.com>
@requie requie changed the title Add CVSS v3.1 vector strings to all benchmark severity objects feat: pattern-specific definitions and severity vectors for all 200 records Sep 30, 2026
@requie
requie merged commit 9360ad2 into main Sep 30, 2026
8 checks passed
@requie
requie deleted the feat/pattern-definitions-and-severity branch September 30, 2026 20:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant