Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "adept",
"version": "4.6.2",
"version": "4.7.0",
"description": "Development-workflow skills: design, TDD, adversarial review, shipping, and campaign orchestration for Claude Code and Codex.",
"author": {
"name": "David Christensen"
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
# 0057 — Review grades reachability against a frozen failure model

## Status

Accepted (2026-09-09)

## Context

`$gauntlet` is instructed to break confidence in a target, and its document attack list names
empty, large, malformed, concurrent, and hostile cases as things to look for. Nothing tells it
which of those matter for the change in front of it. A design under `$spellcraft` states what
can go wrong only when the change is security-relevant — the threat model at
`skills/spellcraft/SKILL.md` — so an ordinary correctness or robustness change reaches review
with no stated actors, deployments, or accepted failures at all. The reviewer supplies its own,
and it supplies the worst case: every deployment, every actor, every input.

The observed shape is a spec that says "all" and a review that answers with an open series of
"what about this?" findings, one per instance the word would have to cover. `$trial-loop`'s
`rejected-with-evidence` disposition already names the way out — a finding that "presumes a
requirement or threat model nothing claims" — but on a non-security change no such model was
ever written, so the author refutes each instance from scratch. The reviewer generates
against an unbounded quantifier for free; the author pays per finding.

ADRs 0049 through 0053 bounded what a pass costs and when a loop stops. None of them bounds
what a pass reports, and the finding bar's per-finding questions cannot: they raise the quality
of each finding without touching the count.

## Decision

**1. Every spec carries a `Failure model` section, on both lanes.** Four entries: actors and
deployments, invariants and assets at stake, accepted failure classes each with its reason, and
classes covered elsewhere with their owner. In the light lane it is a third-level subsection of
`Scope` and counts against the caps; in the full lane it is its own section. It is sized to
the change. The security threat model becomes its security-specific extension and keeps its
four items.

**2. The model is a review target during the design review and frozen after it.** The design
reviewer challenges each entry once on the merits. After the review closes, `$quest` names the
section to the branch reviewer on a `failure model:` line of the review block, carried by a
new optional `$trial-loop` input. The line is not a ninth charter field: it carries no scope
authority and is not hashed.

**3. `$gauntlet` grades reachability against the model.** A trigger inside a named
deployment or against a named invariant is a finding on ordinary terms. A class the model
accepts is dropped and disclosed as a suppression, on the same terms as governing-ADR
re-litigation, with the entry named. A trigger outside the named deployments is at most a
`medium` note stating the gap, reported once. A defensible case that the model itself is wrong
is one finding against the entry, with evidence, at whatever severity the evidence supports.

**4. A universal claim is one finding.** "All", "every", "any", "never", "always", or an
unbounded "must" with no stated closure earns exactly one finding against the quantifier,
whose remedy is to bound or ground it. Its instances are not enumerated. `$spellcraft`'s
self-review bounds every such word before the reviewer sees it.

**5. `$trial-loop` cites the model to reject in one line.** A finding the frozen model accepts
or places outside the named deployments takes `rejected-with-evidence` by citing the entry. A
defensible finding against an entry takes `blocked`, because changing a frozen model is a
design decision. Adding an accepted class mid-cycle to retire a finding is the exclusion-gaming
the charter section already forbids.

## Consequences

- A reviewer has a stated referent for "does this matter here", written by the author before
the review and challenged once. Instance enumeration against an accepted class or an unnamed
deployment stops being reportable as findings.
- The author owes one line per rejected finding instead of a fresh refutation each, which is
the cheap re-disposition the iteration cap depends on.
- Every spec grows by a few lines, including a light spec inside its 60-line cap. On a small
change that is one operator, one invariant, and one or two accepted classes.
- A wrong model is now a way to under-review: an accepted class nobody should have accepted
suppresses real findings. Three things hold that: the design review attacks each entry, the
branch reviewer may attack an entry with evidence at blocking severity, and every acceptance
is disclosed in `suppressions` and surfaced on `approve`.
- `suppressions` entries now carry either `adr` or `failure_model`. Readers that assumed every
entry names an ADR see an entry that names something else; the `suppressed_count` contract is
unchanged.
- A `no-spec` run passes `none` and reviews exactly as before. Standalone `$gauntlet` on a
document with no model gets only the quantifier rule.

## Considered & rejected

- **Add the failure model as a ninth charter field.** verified: at `9d866fc`,
`rg --no-config -n 'eight' skills/quest/SKILL.md skills/trial-loop/SKILL.md
skills/spellcraft/SKILL.md` matches ten lines stating the eight-field contract, and
`skills/spellcraft/SKILL.md` freezes it before design with "Use a complete caller-supplied
charter unchanged" — the charter is fixed at `$quest`'s scope checkpoint, before the code is
read, so a field that needs the codebase cannot be written there. judgment: the model also
needs to be attackable, and charter fields are authority, not claims.
- **Cap the number of findings per pass.** judgment: a count cap drops the last real finding
as readily as the last nitpick. The CriticGPT result (McAleese et al., 2024) is that catches
and nitpicks rise together with claim count and the only lever that separated them was a
learned precision term; a raw count is not that lever.
- **Keep the threat model conditional and rely on `rejected-with-evidence` (do nothing).**
verified: at `9d866fc`, `skills/spellcraft/SKILL.md` gates the threat model on the `$quest`
step 6 security triggers, and `skills/trial-loop/SKILL.md` names "a threat model nothing
claims" as rejection evidence — a referent that, on a non-security change, nothing writes.
- **Let the reviewer derive the model itself.** judgment: a model the reviewer writes is graded
against a premise the author never saw and cannot cite, and it is rewritten on every pass.
The value of the section is that it is authored once, before review, and challenged once.
- **Treat an outside-the-model trigger as a suppression rather than a note.** judgment: a
suppression is visible only by count and entry, while a note keeps the gap stated in the
findings where the author decides whether the model was incomplete. The note costs one
disposition; the gap it names is the information.
3 changes: 2 additions & 1 deletion references/review-depth.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,7 +81,8 @@ Then apply `$trial-loop` step 2's checks, which a single pass does not get for f
- treat an `approve` carrying a non-zero `blocking_count`, or a `blocking_count` above
`findings_count`, as malformed: rerun once, then stop as blocked;
- open the artifact when `findings_count > 0` **or** `suppressed_count > 0`, and
surface each suppression (concern plus ADR) in the transcript.
surface each suppression (concern plus the ADR or failure-model entry that settled it)
in the transcript.

Validate the full artifact before routing it, exactly as `$trial-loop` step 2
does. Every finding has exactly one recognized `surface` (`in | adjacent`) and
Expand Down
78 changes: 65 additions & 13 deletions skills/gauntlet/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,10 @@ and emphasis characters, and ignoring trailing punctuation or emphasis. Matching
case-sensitive; the start of the supplied invocation text counts as a line start. So `CHARTER`, ` CHARTER (…)`,
`CHARTER:`, `- CHARTER:` and `**CHARTER:**` all match. That line and every token after it is
**focus text** — never a target, never a flag. A path named there is prose describing a permitted
change surface, not a file to review. Classify only the tokens *before* that line with the rules
below.
change surface, not a file to review. The one line you *read* from is `failure model:`, which
names a document section to consult under *Failure model* below — read it, never review it as a
target unless it is also among the target tokens. Classify only the tokens *before* that line
with the rules below.

**When a `CHARTER` label was found** and classifying the tokens before it leaves no target token
and no `--base`/`--working-tree` flag, do **not** fall through to the working-tree default. A
Expand Down Expand Up @@ -142,6 +144,49 @@ it to what the charter supports. Recommend controls, transactions, persistence,
or other machinery only when the frozen charter or an explicit user decision authorizes
the guarantee.

**A universal claim is one finding.** "All", "every", "any", "never", "always", or an
unbounded "must" with no stated closure is an ungrounded guarantee of the kind above, and it
earns exactly one finding — against the quantifier — whose remedy is to bound the claim to a
named set or ground it in the charter. Its instances are not findings: do not enumerate the
cases the word would have to cover, because each is the same defect restated and the list has
no end. Once the claim is bounded, a case inside the bound that the target mishandles is a
finding on its own terms.

### Failure model — the frozen answer to what can go wrong

A design written under `$spellcraft` carries a `Failure model` section: the actors and
deployments the change serves, the invariants and assets it must not break, the failure
classes it accepts with a reason each, and the ones another owner covers. It is the target's
answer to "what can go wrong here that matters", frozen when the design review closed. A
caller names it with a `failure model:` line in the `CHARTER` block; a design review meets it
inside the target. Read it before forming findings. `none`, or no line at all → skip this
section silently; nothing below applies.

Grade reachability against the model, not against the worst deployment you can imagine:

- **Inside the model.** A trigger that uses a named actor, deployment, or input, or that
breaks a named invariant, is a finding on its ordinary terms.
- **Accepted by the model.** A would-be finding whose failure class the model accepts, with a
reason that holds, is not reported as an instance. Drop it and disclose it: a `suppressions`
entry naming the concern and the entry that accepted it, exactly as a governing-ADR drop is
disclosed, on an `approve` too.
- **Outside the model.** A trigger that needs an actor, deployment, or condition the model
neither names nor accepts is at most a note — `medium`, with the model gap stated in the
body — because under the model's own premise it is unreachable. It takes one disposition
downstream, and a second such note adds nothing the first did not say: report the gap once.
- **Against the model.** When the model itself is wrong — an accepted class is one the
charter's outcome or completion criteria require, a reason does not hold, or a named
deployment is contradicted by the repository's own evidence (a CI workflow, a published entry
point, a documented consumer) — raise **one** finding against that entry with the evidence.
It takes the severity its evidence supports and may block. This is the counterpart of a
supersession proposal: the way past the model runs through the entry, never around it.

The model is a review target whenever it sits inside the artifact under review. Challenge
each entry once, on the merits, as an entry; do not also report the instances an entry you
reject would have covered — the finding against the entry carries them. The model bounds what
you report as instances. It never lowers the bar on anything it names, and it cannot make a
correctness dependency of the charter's outcome disappear by accepting it.

### Governing ADRs — respect accepted decisions

Some repos record architectural decisions as ADRs (typically `docs/adr/`). An
Expand Down Expand Up @@ -188,9 +233,10 @@ governing ADRs and respect them — without letting that silence genuine new ris
- **Disclose every suppression.** When you suppress a would-be finding as
governing-ADR re-litigation, record it: add an entry to the `suppressions` array
(`--json`) naming the concern you dropped and the ADR that settled it — or, in
markdown mode, a **Suppressed (governing ADR)** block (see Output). Do this **even
when the verdict is `approve`**, since that is the case a caller cannot infer from
the verdict.
markdown mode, a **Suppressed** block (see Output). A finding dropped because the
failure model accepts its class is disclosed the same way, naming the entry instead
of an ADR. Do this **even when the verdict is `approve`**, since that is the case a
caller cannot infer from the verdict.
Over-suppression is the main hazard of this stance, so silent suppression is not
allowed — a suppressed finding must leave an auditable trace, exactly as the
budget-exhaustion escape does.
Expand Down Expand Up @@ -298,15 +344,17 @@ For every finding use real line numbers from the file you read or the diff hunk.
- <action>
- <action>

**Suppressed (governing ADR):**
**Suppressed:**
- <concern you dropped> — settled by ADR <NNNN>
- <concern you dropped> — accepted by failure model entry <entry>
```

Include the **Suppressed (governing ADR)** block whenever you dropped a finding as
governing-ADR re-litigation — it is the markdown counterpart of the `suppressions`
array and **persists even on an `approve` verdict** — whichever finding sections that
verdict carries — because an approve that suppressed a real finding is the case a
reader most needs to see. Omit the block only when nothing was suppressed.
Include the **Suppressed** block whenever you dropped a finding as governing-ADR
re-litigation or as a class the failure model accepts — it is the markdown counterpart
of the `suppressions` array and **persists even on an `approve` verdict** — whichever
finding sections that verdict carries — because an approve that suppressed a real
finding is the case a reader most needs to see. Omit the block only when nothing was
suppressed.

Omit either section when it is empty. An `approve` has no **Findings (blocking)**
section by definition, but it keeps its **Notes (non-blocking)** section whenever notes
Expand Down Expand Up @@ -339,11 +387,15 @@ When `--json` is present, the skill's **output artifact** is exactly this JSON o
],
"next_steps": ["..."],
"suppressions": [
{ "concern": "one line: the finding you dropped", "adr": "0002" }
{ "concern": "one line: the finding you dropped", "adr": "0002" },
{ "concern": "one line: the finding you dropped", "failure_model": "<the entry that accepted it>" }
]
}
```

A suppression carries `concern` and exactly one of `adr` or `failure_model` — what settled
it. An entry with neither names nothing a caller can audit and is malformed.

`verdict` is `approve` when no **blocking** (`critical` or `high`) finding exists, and `needs-attention` otherwise. `findings`, `next_steps`, and `suppressions` may be empty arrays but must be present.

**One `findings` array, not two.** Blocking findings and notes live in the same array and are told apart by `severity` — the markdown rendering splits them into two sections, the JSON does not. A second array would be a second place for a severity to be recorded, free to disagree with the first.
Expand All @@ -353,7 +405,7 @@ is emitted **first**, per *Finding bar*. A consumer may read it; its job is done
before any consumer sees it. `surface` routes the finding after review and never
changes its severity or whether it contributes to `blocking_count`.

`suppressions` is the machine-readable form of the "Disclose every suppression" rule — populate it whenever you drop a finding as governing-ADR re-litigation, **even when the verdict is `approve`** (that is exactly the case a caller cannot see from the verdict alone).
`suppressions` is the machine-readable form of the "Disclose every suppression" rule — populate it whenever you drop a finding as governing-ADR re-litigation or as a class the failure model accepts, **even when the verdict is `approve`** (that is exactly the case a caller cannot see from the verdict alone).

### Severity vocabulary

Expand Down
6 changes: 4 additions & 2 deletions skills/quest/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -545,8 +545,10 @@ below retains the `detect-evil` route and its `security` lens.

On `iterating`, run `$trial-loop --reviewer gauntlet --base <BASE_BRANCH> <composed focus>`. On
`single-pass`, dispatch the one `gauntlet` pass the reference specifies, with the same `--base`
and composed focus, and give each finding its single disposition. Address every defensible
finding and commit after each accepted fix, on either route.
and composed focus, and give each finding its single disposition. On either route the review
block's `failure model:` line names the reviewed spec's `Failure model` section by
repo-relative path and heading — the loop's `failure_model` input — or `none` on a `no-spec`
run. Address every defensible finding and commit after each accepted fix, on either route.

**A blocking finding on a single pass escalates rather than being fixed in place.** Record the
escalation and the finding that caused it, then run the `$trial-loop` invocation above against
Expand Down
Loading