Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,11 +1,11 @@
`report.md` and the coverage producer's Markdown use exactly one canonical
section and preserve array order. Apply public-inline sanitization to code-like
fields: redact secrets and private paths, collapse whitespace, replace
backticks with apostrophes, and use `unavailable` when blank. Apply public-prose
sanitization to `<summary>` and `<exclusion_reason>`: use the same redaction,
whitespace, and fallback rules, then escape backslashes and Markdown code,
emphasis, link, image, heading, and strikethrough delimiters plus HTML angle
brackets. Measured evidence uses:
The coverage producer's Markdown uses exactly one canonical section and
preserves array order; `report.md` does not render it. Apply public-inline
sanitization to code-like fields: redact secrets and private paths, collapse
whitespace, replace backticks with apostrophes, and use `unavailable` when
blank. Apply public-prose sanitization to `<summary>` and `<exclusion_reason>`:
use the same redaction, whitespace, and fallback rules, then escape backslashes
and Markdown code, emphasis, link, image, heading, and strikethrough delimiters
plus HTML angle brackets. Measured evidence uses:

```text
## Scoped coverage evidence
Expand Down
Original file line number Diff line number Diff line change
@@ -1 +1 @@
For every output declared with `Contract: ultrafuzz/findings@2`, read the exact pinned schema named by `Validate against` in the central Ultrafuzz Output Contract and run the exact `Validation command` rendered beside that output. The declared path is authoritative; do not create a differently named findings artifact. That schema alone owns the JSON version, field names, types, enums, required members, and empty form. Producer lanes make only a preliminary severity estimate; later severity review owns final severity. Cite evidence at its real source location, keep base paths free of embedded selectors, and keep independent explanation separate from verbatim or source-location data. Those ownership and evidence semantics remain required in addition to schema validation.
For every output declared with `Contract: ultrafuzz/findings@2`, read the exact pinned schema named by `Validate against` in the central Ultrafuzz Output Contract and run the exact `Validation command` rendered beside that output. The declared path is authoritative; do not create a differently named findings artifact. That schema alone owns the JSON version, field names, types, enums, required members, and empty form. Producer lanes make only a preliminary severity estimate; later severity review owns final severity. Cite evidence at its real source location, keep base paths free of embedded selectors, and keep independent explanation separate from verbatim or source-location data. Those ownership and evidence semantics remain required in addition to schema validation. Producer lanes set the optional `recommendation` only when their cited evidence establishes the fix, as one concise paragraph that names the change with identifiers in backticks and no placeholders, file paths, or links; otherwise they omit it. Deduplication keeps the root finding's `recommendation` byte-for-byte and never adds one, and no later stage adds one.
5 changes: 5 additions & 0 deletions .ultrafuzz/prompts/review/dedupe-findings.md
Original file line number Diff line number Diff line change
Expand Up @@ -199,6 +199,11 @@ several records into one root or family, use the stable union of their canonical
property IDs on the kept record and relevant family variants; do not discard a
property reference during deduplication.

Preserve the root finding's own `recommendation` byte-for-byte on the kept
record, and leave it absent when the root has none. Do not author one, reword
it, or copy one from a duplicate or family variant; the final report renders
only the root's.

Treat each input finding's runtime-normalized `producer_node_id`,
`source_nodes`, and compatibility `source_node_id` as provenance, not agent
commentary. For every retained root, form a stable first-seen union of every
Expand Down
206 changes: 156 additions & 50 deletions .ultrafuzz/prompts/review/final-report.md

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@

### Breaking changes

- **[runtime] [artifacts] [prompts] [cli] [modal] [evmbench] [docs]** The final report's Run summary gains a `Commit` line after `Repository`, taken from the run's saved target identity: `report.json` `run_metadata.target_commit` is now required, a 40- or 64-character lowercase hex commit or `null` when no Git commit was recorded for the evaluated target (rendered as ``- Commit: `none` (no Git commit was recorded for the evaluated target)``), and a report task whose data-governance record is missing or invalid fails with an `artifact-contract` failure. `report.md` no longer renders `Source run ID`, `## Scoped coverage evidence` or `## Artifact validation warnings`: `report.json` keeps `source_run_id`, `source_run_ids`, `coverage_evidence` and `run_metadata.artifact_validation_warnings`, the public bundle now requires both `artifact-validation-warnings.json` and `artifact-validation-warnings.md` whenever the report carries warnings, a fixed sentence after the Run summary says when scoped coverage could not be measured or was measured but incomplete, and the final-report gate now rejects a scoped coverage section or exact-scope coverage score in `report.md` with `REPORT_COVERAGE_EVIDENCE_MARKDOWN_UNEXPECTED` instead of requiring the section (`REPORT_COVERAGE_EVIDENCE_MARKDOWN_MISSING`). Every production issue now ends with `### Remediation`, showing the selected source finding's `recommendation` or a fixed sentence saying that no remediation was recorded; the report stage only carries `recommendation`, so a report row that adds one its source finding lacks now fails verification, and `ultrafuzz/findings@2` producers are asked to set it only when their cited evidence establishes the fix. Backtick spans in finding prose (titles, descriptions, rationales, Proof of Concept steps, family variants and remediation) now render as inline code, and the public projection also redacts a private path written right after a literal `\`, `>`, `&lt;`, a word-starting `|`, or a run such as `/*` or `/<`, so text that now renders more literally cannot reveal one. `Estimated spend` is always a numeric USD estimate from the new `run.json#spend_estimate` (`ultrafuzz.spend-estimate.v1`), which uses recorded costs first, then the model's route-exclusive catalog entry, then the versioned fallback table `ultrafuzz.fallback-pricing.2026-10-01`, and imputes executed attempts that recorded no usage, including the report's own production; `report.md` never shows a `+`, `unavailable` or `partial_pricing`, `run_metadata.estimated_spend` must match `^\$(?:0|[1-9][0-9]*)\.[0-9]{2,10}$`, and `ultrafuzz report` diagnostics now check the spend against `run.json#spend_estimate`: a runtime presentation must equal it, while the agent's report-start figure must only be numeric and is no longer expected to end in `+` or capped by the run's figure; `ultrafuzz stats`, eval accounting and Modal worker results keep accounting v4 with its `+` and `unavailable` labels and partial-pricing semantics, but v4 now looks up a model it has not already priced through the same route-exclusive catalog: one leading `openrouter/` is stripped, an ID containing `/` or starting with `~` is priced only from OpenRouter, `claude-`, `gpt-`, `chatgpt-`, `o<digit>`, DeepSeek and Kimi or Moonshot IDs only from their first-party entries and any other ID not at all, a trailing `[...]` context alias is stripped, and a catalog entry with zero input and output rates is ignored unless the ID ends in `:free`. A model that v4 used to price from whichever aggregator sorted first can therefore stay unresolved in v4, where its events without a recorded cost make the spend partial, while `spend_estimate` prices it at fallback rates; Modal worker results now accept an `available` pricing catalog that reports unresolved models instead of failing validation. A verifier in a restarted controller now compares the report with the run's record of the report-start Run summary, `smithers/final-report-run-metadata/<attempt-id>.json`, instead of deriving it again from the moved `run.json`. Because `ultrafuzz/report@3` changes in place, reports written before this release no longer verify, and no longer render as unchecked reports either, since they fail the schema; finish a run launched on an earlier release with that release. EVMbench grades the copied `report.md` as `audit.md`, so its scores are not comparable across this release. Projects scaffolded earlier keep their old `.ultrafuzz/prompts/review/final-report.md`; delete it and rerun `ultrafuzz init` to pick up the new prompt (#1268).
- **[runtime] [docs]** Without `ULTRAFUZZ_PROVIDER_HOME_ROOT`, the provider-home root is now `~/.ultrafuzz-provider-homes`, which the adapters create with mode `0700`, instead of `$XDG_STATE_HOME/ultrafuzz/provider-homes` (by default `~/.local/state/ultrafuzz/provider-homes`), and `XDG_STATE_HOME` no longer moves it. That root holds the homes of `OpenRouterAgent`, `DeepSeekAgent`, and a `ClaudeAgent`, `CodexAgent` or `KimiAgent` with a `config_dir`. Ubuntu's default umask `0002` makes `~/.local` `0775`, and the provider-home check, which still refuses a group- or world-writable directory above a provider home unless it is sticky, refused the old root there, so those agents could not start (#1236). Nothing is read from or moved out of the old root. `OpenRouterAgent` and `DeepSeekAgent` use API keys and keep no login in their homes, so they only start without their earlier session history. A `ClaudeAgent`, `CodexAgent` or `KimiAgent` with a `config_dir` finds its home empty: with `auth = "subscription"` it is not logged in, and a provider config file kept in the old home, such as a Codex `config.toml` that selects a route, no longer applies. To keep that state, move the old root into place before an agent first runs on this release: `mv "${XDG_STATE_HOME:-$HOME/.local/state}/ultrafuzz/provider-homes" ~/.ultrafuzz-provider-homes`. To log in again instead, first create the home with `(umask 077; mkdir -p ~/.ultrafuzz-provider-homes/<provider>/<config_dir>)`, where `<provider>` is `claude`, `codex` or `kimi`, then log the CLI in with that directory as `CLAUDE_CONFIG_DIR`, `CODEX_HOME` or `KIMI_CODE_HOME`. Codex refuses a home that does not exist, and a plain `mkdir -p` under umask `0002` would make the directories it creates `0775`, which the check refuses, as in #1236. If you set `XDG_STATE_HOME` to keep these homes elsewhere, set `ULTRAFUZZ_PROVIDER_HOME_ROOT` instead. That variable also moves the home of a `ClaudeAgent`, `CodexAgent` or `KimiAgent` without a `config_dir`, from `~/.claude`, `~/.codex` or `~/.kimi-code` (or `CLAUDE_CONFIG_DIR`, `CODEX_HOME` or `KIMI_CODE_HOME`) to `<root>/claude`, `<root>/codex` or `<root>/kimi`. Rerun `ultrafuzz init` to refresh `.smithers/agents/provider-home.ts`. A run launched earlier runs the adapters sealed at its launch, and so keeps the old root, until `resume --refresh-controller`.
- **[config] [runtime] [prompts] [docs]** A property lens that fails no longer stops the campaign. The packaged `default` topology, which `ultrafuzz init` scaffolds, and the packaged `exhaustive` and `invariant-only` topologies change: their `properties` group now has `failure_policy: continue`, and `property-specification-fanin` moves to a new `property-catalog` group, which still halts. The fan-in consolidates the lenses that passed verification, and the strategies, specialists and review run without the failed lens's properties. The report is PARTIAL when the lens ran out of attempts and unverified when it failed its output contract, unless a later `resume --retry-failed` reruns the lens successfully. That resume, which Modal's durable resume always runs, reruns the failed lens, and when the lens's artifact verifier failed, also every node that started after the lens's attempt, which is most of the campaign. In `exhaustive`, `dynamic-strategy-generator` now also runs when some strategy attempts failed, using the ones that succeeded, instead of being skipped. The fan-in, strategies and specialists now have optional inputs, so when the artifact verifier of one of them fails, `state.json` no longer records `output_contracts.missing` or `terminal_disposition: task-output-validation-failure` for it and public eval diagnostics omit its `failure_code`; the verifier's error stays in `last_error`, as it already did for review tasks. The `exhaustive` and `invariant-only` profiles use the new topology after the upgrade. The `default` and `low-cost` profiles run the project's own `.ultrafuzz/topology.yml`, so an existing project keeps the old behaviour there, where one failed lens fails the run before any strategy starts and leaves it without a report, until you make three edits to that file: add `defaults: {failure_policy: continue}` to the `properties` group; add a `property-catalog` group with no `failure_policy` (the scaffold gives it `label: Property catalog` and `color: "#854d0e"`); and change the `property-specification-fanin` node's `group` from `properties` to `property-catalog`. If you have not customized that file, `ultrafuzz topology copy default .ultrafuzz/topology.yml --force` replaces it with the new default instead. Every existing project, whichever profile it runs, should also delete `.ultrafuzz/prompts/properties/property-specification-fanin.md` and rerun `ultrafuzz init`: a project prompt overrides the built-in one under every profile, and only the new prompt tells the fan-in to consolidate the lenses it is given instead of every lens in the topology.
- **[runtime] [docs]** A `failure_policy: continue` group's results are now optional to every node outside that group, including nodes with no `group` and the nodes a dynamic group generates, not only to the group named `review`. Such a node runs without a failed input instead of being skipped, while nodes in the same group still require it. A node that reaches a failed ancestor of its own group only through another group, such as strategy `s2` after specialist `x` after strategy `s1`, now fails its input admission instead of being skipped; that is an `artifact-contract` failure, so the run's report is published unverified. To keep a chain strict, put all of it in one group.
Expand Down
6 changes: 4 additions & 2 deletions docs/config.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,8 +210,10 @@ what it does not know. An unreadable wire, including inherited history torn by
a killed attempt, leaves that invocation's usage absent instead of failing the
invocation. Kimi model pricing resolves against the Moonshot
provider entry in the pricing catalog, so the configured alias must match a
Moonshot catalog model id such as `kimi-k3`; anything else is reported as an
unresolved model instead of being priced from a same-named third-party entry.
Moonshot catalog model id such as `kimi-k3`. Anything else stays an unresolved
model in accounting v4 instead of being priced from a same-named third-party
entry, and the report's spend estimate prices it at the documented
[fallback rates](reference/artifacts-reports.md#spend-estimate-method).
Subscription runs are not billed per token, so the published cost is an
API-comparison estimate at Moonshot list rates.

Expand Down
6 changes: 6 additions & 0 deletions docs/how-to/review-findings.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,5 +95,11 @@ For each finding you might act on:
4. Keep protocol-specific edits separate from the raw generated artifact so the
review trail stays clear.

Each production issue in `report.md` ends with `### Remediation`. It shows the
`recommendation` the finding's producer recorded, copied unchanged, or a fixed
sentence saying that none was recorded; Ultrafuzz never infers one. Treat a
recorded remediation as a starting point for your own fix, not as a reviewed
patch.

Only materialize files after review, and treat copied files as ordinary
unstaged working-tree changes.
11 changes: 8 additions & 3 deletions docs/how-to/run-evals-on-modal.md
Original file line number Diff line number Diff line change
Expand Up @@ -467,9 +467,14 @@ generation, launch generation and attempt, whether model work started, node
counts, checkpoint age and digest, exit category, runtime, aggregate usage,
pricing provenance, and a generic diagnostic code. They never contain
source text, prompts, findings, provider output, exception text, or raw
artifacts. By default, `collect` copies `status.json`, `result.json`, the generic
worker lifecycle log, and an allowlisted `public-eval-diagnostics.json` when a
public worker reached the post-eval gate. It also writes
artifacts. Pricing provenance restates the run's accounting v4 catalog source,
status, and resolved and unresolved model counts. A status of `available` can
still count unresolved models, which have no pricing route or no usable price
on it; their events without a recorded cost count in
`usage.unpriced_event_count`, and `usage.partial_pricing` is then true. By default,
`collect` copies `status.json`, `result.json`, the generic worker lifecycle log,
and an allowlisted `public-eval-diagnostics.json` when a public worker reached
the post-eval gate. It also writes
`recovery-lifecycle.json` and a privacy-safe analysis bundle whose recovery
totals are derived from those exact records. The lifecycle projection contains
only typed reasons, timestamps, aggregate node counts, fingerprints, and
Expand Down
Loading
Loading