From a01cf6e16f7509bd573ca7032868b81cfd8bf472 Mon Sep 17 00:00:00 2001 From: David Christensen Date: Wed, 9 Sep 2026 09:30:11 -0700 Subject: [PATCH] feat: bound review by a frozen failure model and a quantifier rule Every spec now carries a Failure model section on both lanes: actors and deployments, invariants at stake, accepted failure classes with reasons, and classes covered elsewhere. The security threat model becomes its extension. The design review challenges each entry once; afterwards $quest names the section to the branch reviewer on a `failure model:` line of the review block, carried by a new optional $trial-loop input. $gauntlet grades reachability against the model: accepted classes are disclosed suppressions, triggers outside the named deployments are at most a note reported once, and a wrong entry is one finding against it. A universal claim with no closure is one finding against the quantifier, never an enumeration of its instances. $trial-loop cites a model entry to reject a finding in one line and treats a defensible finding against an entry as blocked. Records the decision as ADR 0057 and bumps the plugin to 4.7.0. --- .claude-plugin/plugin.json | 2 +- ...hability-against-a-frozen-failure-model.md | 103 ++++++++++++++++++ references/review-depth.md | 3 +- skills/gauntlet/SKILL.md | 78 ++++++++++--- skills/quest/SKILL.md | 6 +- skills/spellcraft/SKILL.md | 75 ++++++++++++- skills/trial-loop/SKILL.md | 46 +++++--- 7 files changed, 276 insertions(+), 37 deletions(-) create mode 100644 docs/adr/0057-review-grades-reachability-against-a-frozen-failure-model.md diff --git a/.claude-plugin/plugin.json b/.claude-plugin/plugin.json index 21d9dbd..ef94181 100644 --- a/.claude-plugin/plugin.json +++ b/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "adept", - "version": "4.6.2", + "version": "4.7.0", "description": "Development-workflow skills: design, TDD, adversarial review, shipping, and campaign orchestration for Claude Code and Codex.", "author": { "name": "David Christensen" diff --git a/docs/adr/0057-review-grades-reachability-against-a-frozen-failure-model.md b/docs/adr/0057-review-grades-reachability-against-a-frozen-failure-model.md new file mode 100644 index 0000000..9e50c04 --- /dev/null +++ b/docs/adr/0057-review-grades-reachability-against-a-frozen-failure-model.md @@ -0,0 +1,103 @@ +# 0057 — Review grades reachability against a frozen failure model + +## Status + +Accepted (2026-09-09) + +## Context + +`$gauntlet` is instructed to break confidence in a target, and its document attack list names +empty, large, malformed, concurrent, and hostile cases as things to look for. Nothing tells it +which of those matter for the change in front of it. A design under `$spellcraft` states what +can go wrong only when the change is security-relevant — the threat model at +`skills/spellcraft/SKILL.md` — so an ordinary correctness or robustness change reaches review +with no stated actors, deployments, or accepted failures at all. The reviewer supplies its own, +and it supplies the worst case: every deployment, every actor, every input. + +The observed shape is a spec that says "all" and a review that answers with an open series of +"what about this?" findings, one per instance the word would have to cover. `$trial-loop`'s +`rejected-with-evidence` disposition already names the way out — a finding that "presumes a +requirement or threat model nothing claims" — but on a non-security change no such model was +ever written, so the author refutes each instance from scratch. The reviewer generates +against an unbounded quantifier for free; the author pays per finding. + +ADRs 0049 through 0053 bounded what a pass costs and when a loop stops. None of them bounds +what a pass reports, and the finding bar's per-finding questions cannot: they raise the quality +of each finding without touching the count. + +## Decision + +**1. Every spec carries a `Failure model` section, on both lanes.** Four entries: actors and +deployments, invariants and assets at stake, accepted failure classes each with its reason, and +classes covered elsewhere with their owner. In the light lane it is a third-level subsection of +`Scope` and counts against the caps; in the full lane it is its own section. It is sized to +the change. The security threat model becomes its security-specific extension and keeps its +four items. + +**2. The model is a review target during the design review and frozen after it.** The design +reviewer challenges each entry once on the merits. After the review closes, `$quest` names the +section to the branch reviewer on a `failure model:` line of the review block, carried by a +new optional `$trial-loop` input. The line is not a ninth charter field: it carries no scope +authority and is not hashed. + +**3. `$gauntlet` grades reachability against the model.** A trigger inside a named +deployment or against a named invariant is a finding on ordinary terms. A class the model +accepts is dropped and disclosed as a suppression, on the same terms as governing-ADR +re-litigation, with the entry named. A trigger outside the named deployments is at most a +`medium` note stating the gap, reported once. A defensible case that the model itself is wrong +is one finding against the entry, with evidence, at whatever severity the evidence supports. + +**4. A universal claim is one finding.** "All", "every", "any", "never", "always", or an +unbounded "must" with no stated closure earns exactly one finding against the quantifier, +whose remedy is to bound or ground it. Its instances are not enumerated. `$spellcraft`'s +self-review bounds every such word before the reviewer sees it. + +**5. `$trial-loop` cites the model to reject in one line.** A finding the frozen model accepts +or places outside the named deployments takes `rejected-with-evidence` by citing the entry. A +defensible finding against an entry takes `blocked`, because changing a frozen model is a +design decision. Adding an accepted class mid-cycle to retire a finding is the exclusion-gaming +the charter section already forbids. + +## Consequences + +- A reviewer has a stated referent for "does this matter here", written by the author before + the review and challenged once. Instance enumeration against an accepted class or an unnamed + deployment stops being reportable as findings. +- The author owes one line per rejected finding instead of a fresh refutation each, which is + the cheap re-disposition the iteration cap depends on. +- Every spec grows by a few lines, including a light spec inside its 60-line cap. On a small + change that is one operator, one invariant, and one or two accepted classes. +- A wrong model is now a way to under-review: an accepted class nobody should have accepted + suppresses real findings. Three things hold that: the design review attacks each entry, the + branch reviewer may attack an entry with evidence at blocking severity, and every acceptance + is disclosed in `suppressions` and surfaced on `approve`. +- `suppressions` entries now carry either `adr` or `failure_model`. Readers that assumed every + entry names an ADR see an entry that names something else; the `suppressed_count` contract is + unchanged. +- A `no-spec` run passes `none` and reviews exactly as before. Standalone `$gauntlet` on a + document with no model gets only the quantifier rule. + +## Considered & rejected + +- **Add the failure model as a ninth charter field.** verified: at `9d866fc`, + `rg --no-config -n 'eight' skills/quest/SKILL.md skills/trial-loop/SKILL.md + skills/spellcraft/SKILL.md` matches ten lines stating the eight-field contract, and + `skills/spellcraft/SKILL.md` freezes it before design with "Use a complete caller-supplied + charter unchanged" — the charter is fixed at `$quest`'s scope checkpoint, before the code is + read, so a field that needs the codebase cannot be written there. judgment: the model also + needs to be attackable, and charter fields are authority, not claims. +- **Cap the number of findings per pass.** judgment: a count cap drops the last real finding + as readily as the last nitpick. The CriticGPT result (McAleese et al., 2024) is that catches + and nitpicks rise together with claim count and the only lever that separated them was a + learned precision term; a raw count is not that lever. +- **Keep the threat model conditional and rely on `rejected-with-evidence` (do nothing).** + verified: at `9d866fc`, `skills/spellcraft/SKILL.md` gates the threat model on the `$quest` + step 6 security triggers, and `skills/trial-loop/SKILL.md` names "a threat model nothing + claims" as rejection evidence — a referent that, on a non-security change, nothing writes. +- **Let the reviewer derive the model itself.** judgment: a model the reviewer writes is graded + against a premise the author never saw and cannot cite, and it is rewritten on every pass. + The value of the section is that it is authored once, before review, and challenged once. +- **Treat an outside-the-model trigger as a suppression rather than a note.** judgment: a + suppression is visible only by count and entry, while a note keeps the gap stated in the + findings where the author decides whether the model was incomplete. The note costs one + disposition; the gap it names is the information. diff --git a/references/review-depth.md b/references/review-depth.md index ca7fdd4..0969491 100644 --- a/references/review-depth.md +++ b/references/review-depth.md @@ -81,7 +81,8 @@ Then apply `$trial-loop` step 2's checks, which a single pass does not get for f - treat an `approve` carrying a non-zero `blocking_count`, or a `blocking_count` above `findings_count`, as malformed: rerun once, then stop as blocked; - open the artifact when `findings_count > 0` **or** `suppressed_count > 0`, and - surface each suppression (concern plus ADR) in the transcript. + surface each suppression (concern plus the ADR or failure-model entry that settled it) + in the transcript. Validate the full artifact before routing it, exactly as `$trial-loop` step 2 does. Every finding has exactly one recognized `surface` (`in | adjacent`) and diff --git a/skills/gauntlet/SKILL.md b/skills/gauntlet/SKILL.md index 56336ac..c2388ef 100644 --- a/skills/gauntlet/SKILL.md +++ b/skills/gauntlet/SKILL.md @@ -26,8 +26,10 @@ and emphasis characters, and ignoring trailing punctuation or emphasis. Matching case-sensitive; the start of the supplied invocation text counts as a line start. So `CHARTER`, ` CHARTER (…)`, `CHARTER:`, `- CHARTER:` and `**CHARTER:**` all match. That line and every token after it is **focus text** — never a target, never a flag. A path named there is prose describing a permitted -change surface, not a file to review. Classify only the tokens *before* that line with the rules -below. +change surface, not a file to review. The one line you *read* from is `failure model:`, which +names a document section to consult under *Failure model* below — read it, never review it as a +target unless it is also among the target tokens. Classify only the tokens *before* that line +with the rules below. **When a `CHARTER` label was found** and classifying the tokens before it leaves no target token and no `--base`/`--working-tree` flag, do **not** fall through to the working-tree default. A @@ -142,6 +144,49 @@ it to what the charter supports. Recommend controls, transactions, persistence, or other machinery only when the frozen charter or an explicit user decision authorizes the guarantee. +**A universal claim is one finding.** "All", "every", "any", "never", "always", or an +unbounded "must" with no stated closure is an ungrounded guarantee of the kind above, and it +earns exactly one finding — against the quantifier — whose remedy is to bound the claim to a +named set or ground it in the charter. Its instances are not findings: do not enumerate the +cases the word would have to cover, because each is the same defect restated and the list has +no end. Once the claim is bounded, a case inside the bound that the target mishandles is a +finding on its own terms. + +### Failure model — the frozen answer to what can go wrong + +A design written under `$spellcraft` carries a `Failure model` section: the actors and +deployments the change serves, the invariants and assets it must not break, the failure +classes it accepts with a reason each, and the ones another owner covers. It is the target's +answer to "what can go wrong here that matters", frozen when the design review closed. A +caller names it with a `failure model:` line in the `CHARTER` block; a design review meets it +inside the target. Read it before forming findings. `none`, or no line at all → skip this +section silently; nothing below applies. + +Grade reachability against the model, not against the worst deployment you can imagine: + +- **Inside the model.** A trigger that uses a named actor, deployment, or input, or that + breaks a named invariant, is a finding on its ordinary terms. +- **Accepted by the model.** A would-be finding whose failure class the model accepts, with a + reason that holds, is not reported as an instance. Drop it and disclose it: a `suppressions` + entry naming the concern and the entry that accepted it, exactly as a governing-ADR drop is + disclosed, on an `approve` too. +- **Outside the model.** A trigger that needs an actor, deployment, or condition the model + neither names nor accepts is at most a note — `medium`, with the model gap stated in the + body — because under the model's own premise it is unreachable. It takes one disposition + downstream, and a second such note adds nothing the first did not say: report the gap once. +- **Against the model.** When the model itself is wrong — an accepted class is one the + charter's outcome or completion criteria require, a reason does not hold, or a named + deployment is contradicted by the repository's own evidence (a CI workflow, a published entry + point, a documented consumer) — raise **one** finding against that entry with the evidence. + It takes the severity its evidence supports and may block. This is the counterpart of a + supersession proposal: the way past the model runs through the entry, never around it. + +The model is a review target whenever it sits inside the artifact under review. Challenge +each entry once, on the merits, as an entry; do not also report the instances an entry you +reject would have covered — the finding against the entry carries them. The model bounds what +you report as instances. It never lowers the bar on anything it names, and it cannot make a +correctness dependency of the charter's outcome disappear by accepting it. + ### Governing ADRs — respect accepted decisions Some repos record architectural decisions as ADRs (typically `docs/adr/`). An @@ -188,9 +233,10 @@ governing ADRs and respect them — without letting that silence genuine new ris - **Disclose every suppression.** When you suppress a would-be finding as governing-ADR re-litigation, record it: add an entry to the `suppressions` array (`--json`) naming the concern you dropped and the ADR that settled it — or, in - markdown mode, a **Suppressed (governing ADR)** block (see Output). Do this **even - when the verdict is `approve`**, since that is the case a caller cannot infer from - the verdict. + markdown mode, a **Suppressed** block (see Output). A finding dropped because the + failure model accepts its class is disclosed the same way, naming the entry instead + of an ADR. Do this **even when the verdict is `approve`**, since that is the case a + caller cannot infer from the verdict. Over-suppression is the main hazard of this stance, so silent suppression is not allowed — a suppressed finding must leave an auditable trace, exactly as the budget-exhaustion escape does. @@ -298,15 +344,17 @@ For every finding use real line numbers from the file you read or the diff hunk. - - -**Suppressed (governing ADR):** +**Suppressed:** - — settled by ADR +- — accepted by failure model entry ``` -Include the **Suppressed (governing ADR)** block whenever you dropped a finding as -governing-ADR re-litigation — it is the markdown counterpart of the `suppressions` -array and **persists even on an `approve` verdict** — whichever finding sections that -verdict carries — because an approve that suppressed a real finding is the case a -reader most needs to see. Omit the block only when nothing was suppressed. +Include the **Suppressed** block whenever you dropped a finding as governing-ADR +re-litigation or as a class the failure model accepts — it is the markdown counterpart +of the `suppressions` array and **persists even on an `approve` verdict** — whichever +finding sections that verdict carries — because an approve that suppressed a real +finding is the case a reader most needs to see. Omit the block only when nothing was +suppressed. Omit either section when it is empty. An `approve` has no **Findings (blocking)** section by definition, but it keeps its **Notes (non-blocking)** section whenever notes @@ -339,11 +387,15 @@ When `--json` is present, the skill's **output artifact** is exactly this JSON o ], "next_steps": ["..."], "suppressions": [ - { "concern": "one line: the finding you dropped", "adr": "0002" } + { "concern": "one line: the finding you dropped", "adr": "0002" }, + { "concern": "one line: the finding you dropped", "failure_model": "" } ] } ``` +A suppression carries `concern` and exactly one of `adr` or `failure_model` — what settled +it. An entry with neither names nothing a caller can audit and is malformed. + `verdict` is `approve` when no **blocking** (`critical` or `high`) finding exists, and `needs-attention` otherwise. `findings`, `next_steps`, and `suppressions` may be empty arrays but must be present. **One `findings` array, not two.** Blocking findings and notes live in the same array and are told apart by `severity` — the markdown rendering splits them into two sections, the JSON does not. A second array would be a second place for a severity to be recorded, free to disagree with the first. @@ -353,7 +405,7 @@ is emitted **first**, per *Finding bar*. A consumer may read it; its job is done before any consumer sees it. `surface` routes the finding after review and never changes its severity or whether it contributes to `blocking_count`. -`suppressions` is the machine-readable form of the "Disclose every suppression" rule — populate it whenever you drop a finding as governing-ADR re-litigation, **even when the verdict is `approve`** (that is exactly the case a caller cannot see from the verdict alone). +`suppressions` is the machine-readable form of the "Disclose every suppression" rule — populate it whenever you drop a finding as governing-ADR re-litigation or as a class the failure model accepts, **even when the verdict is `approve`** (that is exactly the case a caller cannot see from the verdict alone). ### Severity vocabulary diff --git a/skills/quest/SKILL.md b/skills/quest/SKILL.md index 70fc215..110457b 100644 --- a/skills/quest/SKILL.md +++ b/skills/quest/SKILL.md @@ -545,8 +545,10 @@ below retains the `detect-evil` route and its `security` lens. On `iterating`, run `$trial-loop --reviewer gauntlet --base `. On `single-pass`, dispatch the one `gauntlet` pass the reference specifies, with the same `--base` -and composed focus, and give each finding its single disposition. Address every defensible -finding and commit after each accepted fix, on either route. +and composed focus, and give each finding its single disposition. On either route the review +block's `failure model:` line names the reviewed spec's `Failure model` section by +repo-relative path and heading — the loop's `failure_model` input — or `none` on a `no-spec` +run. Address every defensible finding and commit after each accepted fix, on either route. **A blocking finding on a single pass escalates rather than being fixed in place.** Record the escalation and the finding that caused it, then run the `$trial-loop` invocation above against diff --git a/skills/spellcraft/SKILL.md b/skills/spellcraft/SKILL.md index d8e8409..677f557 100644 --- a/skills/spellcraft/SKILL.md +++ b/skills/spellcraft/SKILL.md @@ -149,8 +149,9 @@ when the change is genuinely small. Skipping it is not. Write or update the design doc under `docs/workflow/specs/`, named `YYYY-MM-DD--design.md`. In the light lane it contains exactly four second-level -sections: `Problem`, `Scope`, `Success`, and `Validation`. It is one independently -implementable unit, not a task breakdown. Its Validation section inventories every material +sections: `Problem`, `Scope`, `Success`, and `Validation`, with the failure model required +below as a third-level subsection of `Scope`. It is one independently implementable unit, +not a task breakdown. Its Validation section inventories every material changed contract using the same `focused-test` and `task-test-not-applicable` fields the full plan requires below, including the concrete non-applicability reason rather than a prose test. One page means no more than 500 words and 60 physical lines, including headings and blank lines; @@ -230,8 +231,15 @@ what you find inline: it still need decomposing? - **Two-way ambiguity** — any requirement a competent reader could take two ways. Settle it and say which reading the spec means. -- **Light-spec completeness** — when routed light, verify the exact four-section shape, map every - success criterion to a supported Validation entry, and recheck the 500-word and 60-line caps. +- **Failure model** — the section exists with its four entries, sized to the change, and no + accepted class is one a completion criterion requires. +- **Universal words** — every "all", "every", "any", "never", and "always" in `Success`, a + completion criterion, or a guarantee is bounded to a named set, and what falls outside the + set is accepted or covered in the failure model. The reviewer raises an unbounded one as a + finding against the word; settle it here instead. +- **Light-spec completeness** — when routed light, verify the exact four-section shape with the + `Failure model` subsection under `Scope`, map every success criterion to a supported + Validation entry, and recheck the 500-word and 60-line caps. This pass is cheap and catches the defects an adversarial review would otherwise spend an iteration discovering. It does not replace step 3 — and in the full lane, with the plan @@ -313,7 +321,46 @@ Add to the spec: The eval cases are acceptance criteria: `$forge` implements them as executable tests when their observable contract supports one, or as the spec's bounded evaluation when it does not. -### Security-relevant changes require a threat model +### Every spec carries a failure model + +Answer, before anyone reviews the design, what can go wrong here that matters. Every spec +carries a section headed `Failure model` — its own second-level section in the full lane, a +third-level subsection of `Scope` in the light lane, where it counts against the caps — with +four entries. Each entry is a list of short lines, not prose; an entry with nothing in it says +`none`. + +1. **Actors and deployments** — who runs or calls the changed code and where: a local operator + at a terminal, a CI job, an authenticated tenant, another service, anonymous traffic. Name + the deployments the change is designed for. An unnamed one is outside the model, and the + reviewer grades a trigger that needs it as a note, not a blocker. +2. **Invariants and assets at stake** — what is expensive to get wrong: data that cannot be + recovered, state another actor reads, money, availability, a published contract. This is + where the assessment's hazards land, and where a reviewer's blocking findings come from. +3. **Accepted failure classes** — each with the reason it is accepted: not reachable in the + named deployments, tolerated because its cost is bounded and stated here, or already held by + a named existing guardrail. An acceptance the charter's outcome or completion criteria + contradict is a defect the review will find; never write one to make a criterion cheaper. +4. **Covered elsewhere** — failure classes another owner holds, with the owner: a record, an + issue, a guardrail. + +Size it to the change. A small change with no hazards gets a few lines — one named operator, +one invariant, one or two accepted classes. Silence reads as coverage: a class the model neither +names nor accepts is one the reviewer may still raise, so the section's value is in what it +declines, stated. It is a review target — the design review challenges each entry once, on the +merits — and once that review closes it is frozen with the design. The branch reviewer then +receives it as the frozen answer to what matters: instances of an accepted class are disclosed +suppressions rather than findings, a trigger outside the named deployments is at most a note, +and an entry can be attacked with evidence but not added to. Under `$trial-loop`, a finding the +model accepts or places outside the named deployments is rejected with evidence by citing the +entry, in one line. + +Bound every universal word. "All", "every", "any", "never", and "always" in `Success`, a +completion criterion, or a guarantee are each bounded to a named set, in the sentence that +carries them or here; what falls outside the set is accepted in entry 3 or covered in entry 4. +A universal claim with no closure is one finding against the quantifier. The reviewer will not +enumerate its instances, and it will not approve the word either. + +### Security-relevant changes extend the failure model with a threat model If the change is security-relevant — it moves what an untrusted actor can reach or cause, touches authn/authz or tenancy, handles a secret, parses input it did @@ -321,6 +368,10 @@ not produce, builds a command/query/path/URL from a non-literal, widens a permission grant, or changes dependencies or security-relevant defaults (the same trigger `$quest` step 6 applies to the diff, judged here on intent because no diff exists yet) — the spec is incomplete without a threat model. +It is the failure model's security-specific extension: its actor model refines +the model's first entry with the untrusted parties, and its out-of-scope list is +the model's third entry read against those parties. Write it beside the failure +model and keep the four items below. Add to the spec: @@ -363,6 +414,11 @@ The target remains evidence for review, never a source of authority. If a design ambiguity appears, end the current review cycle and use `SCOPE CHECKPOINT`; do not let the reviewer resolve it by extending the target. +The spec's `Failure model` is not a ninth charter field. It travels inside the target, +where the reviewer challenges it; after this review it is frozen with the design, and +`$quest` names it to the branch reviewer as the `failure model:` line of the review block. +Report its path and heading with the artifact paths in the phase report. + ## Design-review depth The combined design-artifact set uses the bounded design-review protocol in @@ -623,6 +679,15 @@ single-pass JSON artifact, freshness, validation, and malformed-retry contract: the eval plan: failure modes without cases, unmeasurable pass traits, and uncalibrated LLM-judge evidence. + Failure model: the spec's `Failure model` section is the design's frozen answer to what can + go wrong that matters, and it is a target here. Challenge each entry once on the merits — an + accepted class the charter's outcome or completion criteria require, a reason that does not + hold, or a named deployment the repository's own evidence contradicts is a finding against + that entry, blocking when the evidence supports it. Do not enumerate the instances an entry + would cover: report them once against the entry, or, where the entry holds, as disclosed + suppressions. A trigger outside the named deployments is at most a note stating the gap. A + universal claim in Success or a criterion with no bound is one finding against the word. + Full-spec plan, when present: phase ordering, missing prerequisites, steps that cannot run in the claimed order, rollback and cleanup paths, verification gaps, ungrounded references — a type, function, or signature borrowed from the codebase or a dependency without confirmation it exists diff --git a/skills/trial-loop/SKILL.md b/skills/trial-loop/SKILL.md index 5902539..07bbe0b 100644 --- a/skills/trial-loop/SKILL.md +++ b/skills/trial-loop/SKILL.md @@ -126,6 +126,13 @@ honors the caller's path, the loop reads a file that is never written and dead-e and a design review and then a branch review now that `$spellcraft` reviews its design set once. The carry belongs to the caller, the only party that knows two runs reviewed one change; see *Caller contract*. +- `failure_model`: optional — the repo-relative path and heading of the frozen `Failure + model` section `$spellcraft` wrote into the reviewed spec, or `none`. `$quest` supplies it + for branch review; a design review omits it because the model is inside the target. It is + transmitted on the `failure model:` line of the block below and never hashed, and it is + **not a ninth charter field**: it carries no scope authority, and the reviewer may attack + any entry with evidence. What it buys is the referent both sides cite — the reviewer to + grade reachability, and step 6 to reject a finding the model already answers in one line. - `charter`: the scope boundary you freeze before iteration 1 (below). Not an argument the caller types — you derive it. @@ -223,16 +230,18 @@ deferral consumes the whole budget by itself. Watch the other branch too. If the reviewer *does* honor an exclusion and drops a finding, that drop is invisible: `suppressions` and `suppressed_count` cover -governing-ADR re-litigation only, so a charter-driven drop is counted nowhere and the -loop cannot audit it. Treat a finding that stops recurring as unproven, not resolved. +governing-ADR re-litigation and failure-model acceptances only, so a charter-driven drop +is counted nowhere and the loop cannot audit it. Treat a finding that stops recurring as +unproven, not resolved. **A material charter change ends the cycle.** Do not add an exclusion after a finding in order to obtain `approve`. If remediation would materially expand or alter the outcome, completion criteria, a public contract, the persistence model, the -threat model, or the permitted surface, stop, get the authority, update the -charter, and start a **new** cycle with the iteration count reset — an -out-of-charter fix smuggled into iteration 3 is the failure this rule exists to -catch. +threat model, the frozen failure model, or the permitted surface, stop, get the +authority, update the charter, and start a **new** cycle with the iteration count +reset — an out-of-charter fix smuggled into iteration 3 is the failure this rule +exists to catch. An accepted failure class added mid-cycle to retire a finding is +the same gaming as an exclusion added for it. The reset is bounded and visible, or it is just a longer cap. Name in the report who authorized each charter change and what changed, carry every prior cycle's @@ -253,6 +262,7 @@ provenance: exclusions: surface: ambiguities: +failure model: focus: Repeat up to `iteration_budget` iterations (2 unless the caller raised it, 3 without @@ -503,9 +513,9 @@ worker. Do not use step 2's malformed-return retry to replace a worker whose end interactive scope checkpoint, or use the existing unattended park path, only when the unresolved finding itself needs a design decision or authority. 4. If `verdict` is `approve`: when `suppressed_count > 0`, surface each `suppressions` - entry (concern + ADR) in the transcript — an `approve` that suppressed a - governing-ADR finding is exactly the over-suppression case the verdict alone hides, - so it must not advance invisibly. The exit-disclosure rule under *Stop conditions* + entry (concern + the ADR or failure-model entry that settled it) in the transcript — + an `approve` that suppressed a finding is exactly the over-suppression case the + verdict alone hides, so it must not advance invisibly. The exit-disclosure rule under *Stop conditions* also applies here, as it does on every exit. Before exiting, route each `surface: adjacent` note to the `follow-up-candidate` disposition defined in step 6 and disclose its public-safe table row. Do not edit it. Then exit the @@ -544,12 +554,17 @@ worker. Do not use step 2's malformed-return retry to replace a worker whose end disposition, not the residual gap — a change that aggravates a pre-existing defect fixes its own contribution under `accepted-fixed` and defers the rest here, stating the non-regression boundary; - - `rejected-with-evidence` — unsupported, or it presumes a requirement or - threat model nothing claims. The evidence is what makes it a disposition rather - than a dismissal: name what the finding assumes and what refutes it. "It is only - about wording" is not evidence, and neither is the cost of the fix; or + - `rejected-with-evidence` — unsupported, or it presumes a requirement, deployment, + or threat nothing claims. The evidence is what makes it a disposition rather + than a dismissal: name what the finding assumes and what refutes it. A frozen + failure-model entry is that evidence when it accepts the finding's class or places + its trigger outside the named deployments — cite the entry, in one line, and move + on. "It is only about wording" is not evidence, and neither is the cost of the fix; + or - `blocked` — required for correctness, but needs authority, a design - decision, or a material charter expansion. + decision, or a material charter expansion. A defensible finding *against* a frozen + failure-model entry lands here: the model was settled by the design review, so + changing it is a design decision, not a fix. **Resolve at the size of the risk.** Every finding gets a disposition; what scales is what the disposition costs. Where the smallest honest fix would add more to the target @@ -688,7 +703,8 @@ resolves its target from a `git status` the commit just emptied, reviews nothing returns `approve` with a fresh artifact and a matching `run_id`. That is the least supervised path in the whole loop. -Then disclose every suppression (concern + ADR), every `follow-up-candidate` +Then disclose every suppression (concern + the ADR or failure-model entry that settled +it), every `follow-up-candidate` (its public-safe table row), every `deferred-tracked` concern (concern + owning record path or tracker issue), and every `rejected-with-evidence` finding (concern + the pass that raised it) recorded anywhere in the **run**, across all