Skip to content

feat(report)!: target commit, remediation, simplified Markdown, and an always-numeric spend estimate (#1268) - #1276

Open
aviggiano wants to merge 9 commits into
unstablefrom
fix/1268-final-report-commit-remediation-spend
Open

aviggiano wants to merge 9 commits into
unstablefrom
fix/1268-final-report-commit-remediation-spend

Conversation

@aviggiano

@aviggiano aviggiano commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements #1268. The final report now names the evaluated commit, drops low-level diagnostics from report.md, renders remediation for every production finding with readable inline code, and always shows a numeric Estimated spend. The prompt, report contract and projection, renderer, gates, docs, and regression tests change together.

  1. Commit provenance. The Source run ID row is replaced by - Commit: <full hash>, read from the sealed data-governance record (target.commit). report.json run_metadata.target_commit is required: a 40/64-hex string, or null when no Git commit was recorded (rendered as - Commit: `none` (no Git commit was recorded for the evaluated target)). A missing or invalid governance record fails the report task (artifact-contract failure) instead of using a placeholder.

    • Lineage (source_run_id, source_run_ids) and the dirty and worktree digests stay structured-only.
    • The commit is scanned positive-only in the public projection, because the generic 40-hex rule would otherwise redact it.
  2. Simpler Markdown. ## Scoped coverage evidence and ## Artifact validation warnings are no longer rendered.

    • report.json.coverage_evidence and run_metadata.artifact_validation_warnings are unchanged.
    • The coverage producer's own section and the public artifact-validation-warnings.{json,md} companions are byte-identical. The public bundle now requires both companions whenever the report carries warnings.
    • A fixed sentence after the Run summary says when scoped coverage could not be measured, or was measured but incomplete, and the no-issues sentence names unmeasured coverage. PARTIAL, unchecked, campaign, goal, and property-coverage disclosures are unchanged.
    • The final-report gate now rejects the section (REPORT_COVERAGE_EVIDENCE_MARKDOWN_UNEXPECTED) instead of requiring it.
  3. Remediation and readable code.

    • Every production issue ends with ### Remediation, after the Proof of Concept and any family variants. It shows the finding's carried recommendation, or this fixed sentence: "No remediation was recorded for this finding, and Ultrafuzz does not infer one. Confirm the root cause in the description and Proof of Concept before designing a fix."
    • The report stage only carries recommendation. A new semantic-gate error fires when a report row adds one that its source finding lacks. Producers are asked to set it only when their cited evidence establishes the fix, and dedupe keeps the root's.
    • Finding prose now goes through a new findingProse instead of publicProse. Backtick spans render as inline code and intraword _ stays raw. Link, definition, list, heading, HTML, and fence injection stay neutralized.
    • Index anchors now slug the visible title, so they are GitHub-compatible for _ and code spans.
    • publicProse is unchanged for the coverage producer and blocker summaries, whose gates compare bytes.
  4. Always-numeric spend estimate. A new run.json#spend_estimate (ultrafuzz.spend-estimate.v1) is rebuilt on every sync. It prices each attempt from, in order:

    1. the recorded charge;
    2. the actual billing route's catalog rates;
    3. the versioned fallback table ultrafuzz.fallback-pricing.2026-10-01, for missing rates or usage;
    4. imputation, for executed agent attempts that recorded no usage (same-model mean, then run mean, then documented default usage).

    Lineage comes from the source run's persisted estimate. A no-op sync never rewrites run.json.

    • Pricing routes: price lookup is route-exclusive. openrouter/ is stripped; vendor/model and ~ ids are priced only from OpenRouter; claude-, gpt-/o*, deepseek, and kimi ids are priced only from their first-party entries. [1m]-style aliases are stripped, and zero rates are ignored unless the id ends in :free.
    • Report-start snapshot: spend comes from spend_estimate (or the live estimate) plus an imputed estimate of the report's own production. Tokens, spend, and partial_pricing come from one source. This fixes the old loss of partial_pricing: true when cumulative spend was unavailable.
    • Terminal and unchecked presentations restate spend from spend_estimate, which includes report-production usage.
    • A host-only copy of the projection lets a restarted controller's verifier check against what was materialized.
  5. Nothing shows +, unavailable, or partial_pricing in report.md. run_metadata.estimated_spend must match ^\$(?:0|[1-9][0-9]*)\.[0-9]{2,10}$. Method, price sources (catalog provider and model id), fallback rates snapshotted per model, assumption codes, unaccounted attempts, complete, and updated_at live in spend_estimate. ultrafuzz report diagnostics compare runtime presentations to the run.json they were built from and no longer expect +.

Accounting v4 keeps its labels ($x+ and unavailable) and semantics for stats, evals, and Modal. Its catalog lookup for newly fetched models is now route-exclusive. The Modal worker-result gate now accepts available pricing with unresolved models, a documented v4 state that the route change makes more common.

Sample report.md (rendered by this branch's renderer)

## Run summary

- Run ID: `sample-run`
- Repository: `https://github.com/example/vault`
- Commit: `0123456789abcdef0123456789abcdef01234567`
- Elapsed time: `6h 04m`
- Models used: `gpt-5.5`
- Tokens used: `41,203,118`
- Estimated spend: `$38.72`
- Audit profile: `exhaustive`

Scoped coverage could not be measured for this run, so how much of the in-scope code the campaign exercised is unknown.

## [M-01] - `previewRedeem` rounds in the depositor's favour

Depositor can redeem through `previewRedeem` after a fee accrual which leads to `max_withdraw` exceeding the vault's `totalAssets`.

### Severity

- **Impact**: Medium: The vault's accounting drifts by up to one wei per redeem.
- **Likelihood**: High: Any depositor reaches `redeem` through the public flow.

### Proof of Concept

1. Depositor deposits after a fee accrual.
2. Depositor redeems the full share balance.

```solidity
assertLe(vault.maxWithdraw(alice), vault.totalAssets());
```

### Remediation

Round down in `previewRedeem` with `Math.mulDiv(shares, totalAssets(), totalSupply(), Math.Rounding.Floor)`.

Breaking changes

ultrafuzz/report@3 and the run.json schema change in place, following the #1095, #1120, and #1167 precedent (see the CHANGELOG entry):

  • Reports written before this release no longer re-verify. They also no longer render as unchecked reports, because they fail the schema (required target_commit, spend pattern).
  • Finish runs launched on an earlier release with that release.
  • EVMbench grades the copied report.md, so scores across this boundary are not comparable.
  • Scaffolded projects should delete .ultrafuzz/prompts/review/final-report.md and rerun ultrafuzz init.

Design decisions worth reviewing

  • recommendation is carried, never authored, by the report stage. This keeps Remediation traceable to the producer that saw the code. The cost is that most reports will show the fallback sentence until producer lanes adopt the new findings.mdx instruction.
  • Spend imputation and fallback rates are policy. The fallback table is anchored to first-party models.dev list prices fetched 2026-10-02. Default attempt usage (200k uncached + 1.8M cache-read input, 40k output) is the only invented magnitude. Both are versioned and snapshotted, and any use marks the estimate incomplete.
  • The report-start snapshot always has partial_pricing: true, because it always contains an imputed estimate of its own production. The terminal presentation replaces that estimate with recorded usage.
  • Public path redaction is slightly wider. It now also catches a path after a literal \, >, &lt;, a word-starting |, or a run such as /*, so finding text that renders more literally cannot reveal one.

Known limitations / follow-ups (not changed here)

  • A private path directly after a raw < (e.g. cat </srv/x) already published unredacted at base, and still does. It needs its own redaction decision.
  • v4 picks context tiers from the cumulative per-attempt token total (behaviour already at base). The estimate's catalog basis inherits it. Default-usage imputation uses base rates.
  • The in-flight report attempt is imputed at the first model in the agent chain, even when a fallback agent runs it. The terminal restatement corrects this.
  • ULTRAFUZZ_CACHE_READ_RATIO is still undocumented.

Validation

All runs used Node 24.9, locally:

  • Static gates, all passing: pnpm -w format:check, pnpm -w lint, pnpm -w build, pnpm -w knip, node scripts/smithers-patches.mjs --check, ESLINT_PLUGIN_DIFF_COMMIT=5c1a9775 pnpm -w lint:strict:ci, pnpm -w size, pnpm -w docs:check, pnpm -w typecheck.
  • Runtime runtime.test.js (12 shards): 356/357.
    • The one failure, "syncRun aborts or times out a blocked inspection child", hit its 60 s fallback under 12-way parallel load. It passes when run alone (1/1).
  • Runtime supporting suite: 1022/1024.
    • The 2 failures also fail at base 5c1a9775 on this Node version: "a missing or unrenderable prompt fails only its own task at assert-task-inputs" and "generated Smithers verification marker root must be a canonical directory".
  • Bun adapter contract: 53 pass, 1 skip.
  • Other packages: artifacts 336/336, security 23/23, references 22/22. cli 201/201 (all files, incl. e2e campaign-resume), modal 525/525, evals 383/383, config 92/92, prompts 62/62, topology 73/73, evmbench 67/67, dashboard 39/39, scripts/ci 79/79.

Closes #1268.

🤖 Generated with Claude Code

RetriggerConfidence Score: 3/5

The PR is not ready to merge until source-only continuations retain spend lineage and published remediation remains identical to its source recommendation.

Fix All in Claude CodeFindings

  1. P1 Source spend disappears ▶
  2. P1 Rewritten remediation can publish ▶
Fix with agent prompt
### Issue 1
packages/runtime/src/workflow-sync.ts:1493
When a continuation has a source run but no local usage or executed agent attempts, this return skips the source-run estimate. The continuation’s synchronized `run.json` then has no `spend_estimate` or spend lineage, even if the source run has a persisted estimate.

### Issue 2
packages/artifacts/src/semantic-gates.ts:924
If a source finding has a recommendation, this new error check does not catch a report row that changes it. The existing preservation check issues only a warning, so the report can publish different advice under `### Remediation` despite the carry-only contract.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR adds evaluated-commit provenance and a persisted spend estimate, simplifies report Markdown, and renders source-finding recommendations as remediation.

  • Report and run-metadata contracts, projections, pricing, and publication checks change together.
  • The source-only continuation estimate and recommendation-preservation gate need correction.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Usage and attempt ledgers] --> B[Spend estimate]
  S[Source run estimate] --> B
  B --> R[run.json]
  R --> P[Report-start projection]
  F[Source findings] --> G[Preservation gate]
  G --> M[Remediation in report.md]
  P --> M
Loading

Reviews (1) · Last reviewed commit: "docs: describe block-start escaping, tab..."

aviggiano and others added 9 commits October 2, 2026 14:45
Add required run_metadata.target_commit (hex SHA-1/SHA-256 or null) read from
the sealed data-governance record, render it as the Run summary Commit row in
place of Source run ID, keep it through the public projection, and add the
optional run.json spend_estimate contract with the shared
formatEstimatedSpendUsd formatter (#1268).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…and numeric spend

Update the final-report prompt, output-contract templates, prompt structure
tests, reference docs, and CHANGELOG for the Commit row, removed diagnostic
sections, Remediation, readable inline code, and the always-numeric spend
estimate (#1268).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…te (#1268)

Resolve each model's catalog price only from the provider its ID routes to (OpenRouter for gateway IDs, the first-party provider otherwise), reject all-zero rates except :free, strip context aliases for lookup, and return per-model provenance. Add the spend estimator with the dated fallback table, default attempt usage, unaccounted-attempt imputation and source-run lineage; synchronize run.json#spend_estimate with accounting on every pass, including runs whose agent attempts reported no usage; and make the live report-start metrics numeric from the same estimator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…on and readable inline code (#1268)

Stop rendering Scoped coverage evidence and Artifact validation warnings in
report.md (report.json and the public warning companions keep them), state
unmeasured or incomplete scoped coverage in a fixed notice after the Run
summary, end every production issue with Remediation (the carried
recommendation or a fixed fallback), and render finding prose with
findingProse so backtick spans stay inline code while publicProse stays
frozen for every other field. The public projection now redacts a path
written after a literal backslash, a word-starting pipe, or a literal &lt;,
and its Markdown re-scan no longer reads the renderer's own escapes as path
syntax. The final-report gate now rejects a scoped coverage section or
exact-scope score with REPORT_COVERAGE_EVIDENCE_MARKDOWN_UNEXPECTED, a report
row that adds a recommendation its source lacks fails both report loops, and
public bundles require both warning companions when the report carries
warnings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…report

Project the report-start spend from run.json spend_estimate (or the live
estimate), add an imputed estimate for the in-flight report attempt, keep
tokens, spend and partial_pricing from one source, restate spend from
spend_estimate in terminal and unchecked presentations, require the numeric
estimated_spend pattern, persist a host-only copy of the projection for
restarted verifiers, and align ultrafuzz report diagnostics (#1268).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… secret re-scans, and tables (#1268)

- Start-of-block escapes apply only where finding prose opens a block
  (description, Remediation, Proof of Concept steps, disposition entries),
  so a title opening with `[` or a redacted leading path keeps its index link.
- The Markdown path re-scan treats `\*`, `` \` ``, `&lt;`, and `&gt;` after a
  slash, and `&gt;` before one, as the renderer's escapes; `>` is now a path
  boundary in report.json, so a real path after `>` is still redacted.
- The speculative secret pass reads one rendered unit at a time (each line
  split at table pipes and bold delimiters, each fenced block whole) while
  positive detections still read the whole document, so prose ending in
  `token:` no longer fails the public projection.
- Finding titles in the provenance table, prior dispositions, and the
  non-production outcomes table render as finding prose with pipes escaped.
- An unselected reference-expectation property ID holding `<` or `>` renders
  as escaped text instead of raw HTML inside a code span.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ult usage at base rates (#1268)

The spend estimate joined attempts to usage by (workflow run, node,
iteration, attempt), so after `resume --retry-failed` / `--reset-node`
two occurrences sharing an attempt number collapsed into one imputation,
and usage on either one hid the other. It also skipped attempts of a
workflow run that a replay or fork replaced while still pricing that
run's usage, so the estimate dropped and could claim completeness.

- An occurrence is now an attempt-ledger entry (workflow_run_id,
  source_event_sequence). A usage event belongs to the occurrence of its
  workflow run, Smithers task, iteration, and attempt whose start and
  terminal events span its sequence; each occurrence is priced from its
  own latest snapshot and an executed agent occurrence with none is
  imputed, across every workflow run of the run. The Smithers task ID is
  `node:<strategy_attempt_id>`, which every task manifest enforces.
- `spend_estimate.unaccounted_attempts.entries` carry `workflow_run_id`
  and `source_event_sequence` and are unique by them (schema, semantic
  check, builder, docs).
- Default-usage imputation prices catalog models at base rates; the
  2,000,000-token default total no longer selects a context tier.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…v4 route lookup (#1268)

Accounting v4 records pricing_catalog.status "available" when the catalog
was fetched (or no model had a route) yet some models stay unresolved;
route-exclusive pricing makes that common. The Modal worker-result gate
rejected that state, failing status.json/result.json writes and reads.
Drop the rule; uncosted events of an unresolved model still require
usage.partial_pricing through the existing unpriced-event check, and the
unavailable/disabled consistency rules stay.

The #1268 changelog bullet now says v4 keeps its labels and semantics but
looks up newly priced models through the route-exclusive catalog, so some
models an aggregator used to price are now unresolved in v4.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…c path redaction

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@aviggiano
aviggiano requested a review from a team as a code owner October 3, 2026 05:22
const occurrences = spendEstimateAttemptOccurrences(replayNodeAttempts(input.layout).entries, input.usageEntries);
const unaccountedAttempts = occurrences.unaccounted;
const previous = input.metadata.spend_estimate;
if (input.usageEntries.length === 0 && unaccountedAttempts.length === 0) return { changed: previous !== undefined };

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Source spend disappears When a continuation has a source run but no local usage or executed agent attempts, this return skips the source-run estimate. The continuation’s synchronized run.json then has no spend_estimate or spend lineage, even if the source run has a persisted estimate.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/workflow-sync.ts
Line: 1493

Comment:
**Source spend disappears** When a continuation has a source run but no local usage or executed agent attempts, this return skips the source-run estimate. The continuation’s synchronized `run.json` then has no `spend_estimate` or spend lineage, even if the source run has a persisted estimate.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

entry: { row: unknown; path: string },
source: Readonly<Record<string, unknown>>
): SemanticGateIssue[] {
return source.recommendation === undefined && at(entry.row, ["recommendation"]) !== undefined

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Rewritten remediation can publish If a source finding has a recommendation, this new error check does not catch a report row that changes it. The existing preservation check issues only a warning, so the report can publish different advice under ### Remediation despite the carry-only contract.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/artifacts/src/semantic-gates.ts
Line: 924

Comment:
**Rewritten remediation can publish** If a source finding has a recommendation, this new error check does not catch a report row that changes it. The existing preservation check issues only a warning, so the report can publish different advice under `### Remediation` despite the carry-only contract.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

report: include target commit and remediation, simplify Markdown, always estimate spend

1 participant