Conversation
Add required run_metadata.target_commit (hex SHA-1/SHA-256 or null) read from the sealed data-governance record, render it as the Run summary Commit row in place of Source run ID, keep it through the public projection, and add the optional run.json spend_estimate contract with the shared formatEstimatedSpendUsd formatter (#1268). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…and numeric spend Update the final-report prompt, output-contract templates, prompt structure tests, reference docs, and CHANGELOG for the Commit row, removed diagnostic sections, Remediation, readable inline code, and the always-numeric spend estimate (#1268). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…te (#1268) Resolve each model's catalog price only from the provider its ID routes to (OpenRouter for gateway IDs, the first-party provider otherwise), reject all-zero rates except :free, strip context aliases for lookup, and return per-model provenance. Add the spend estimator with the dated fallback table, default attempt usage, unaccounted-attempt imputation and source-run lineage; synchronize run.json#spend_estimate with accounting on every pass, including runs whose agent attempts reported no usage; and make the live report-start metrics numeric from the same estimator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…on and readable inline code (#1268) Stop rendering Scoped coverage evidence and Artifact validation warnings in report.md (report.json and the public warning companions keep them), state unmeasured or incomplete scoped coverage in a fixed notice after the Run summary, end every production issue with Remediation (the carried recommendation or a fixed fallback), and render finding prose with findingProse so backtick spans stay inline code while publicProse stays frozen for every other field. The public projection now redacts a path written after a literal backslash, a word-starting pipe, or a literal <, and its Markdown re-scan no longer reads the renderer's own escapes as path syntax. The final-report gate now rejects a scoped coverage section or exact-scope score with REPORT_COVERAGE_EVIDENCE_MARKDOWN_UNEXPECTED, a report row that adds a recommendation its source lacks fails both report loops, and public bundles require both warning companions when the report carries warnings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…report Project the report-start spend from run.json spend_estimate (or the live estimate), add an imputed estimate for the in-flight report attempt, keep tokens, spend and partial_pricing from one source, restate spend from spend_estimate in terminal and unchecked presentations, require the numeric estimated_spend pattern, persist a host-only copy of the projection for restarted verifiers, and align ultrafuzz report diagnostics (#1268). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… secret re-scans, and tables (#1268) - Start-of-block escapes apply only where finding prose opens a block (description, Remediation, Proof of Concept steps, disposition entries), so a title opening with `[` or a redacted leading path keeps its index link. - The Markdown path re-scan treats `\*`, `` \` ``, `<`, and `>` after a slash, and `>` before one, as the renderer's escapes; `>` is now a path boundary in report.json, so a real path after `>` is still redacted. - The speculative secret pass reads one rendered unit at a time (each line split at table pipes and bold delimiters, each fenced block whole) while positive detections still read the whole document, so prose ending in `token:` no longer fails the public projection. - Finding titles in the provenance table, prior dispositions, and the non-production outcomes table render as finding prose with pipes escaped. - An unselected reference-expectation property ID holding `<` or `>` renders as escaped text instead of raw HTML inside a code span. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ult usage at base rates (#1268) The spend estimate joined attempts to usage by (workflow run, node, iteration, attempt), so after `resume --retry-failed` / `--reset-node` two occurrences sharing an attempt number collapsed into one imputation, and usage on either one hid the other. It also skipped attempts of a workflow run that a replay or fork replaced while still pricing that run's usage, so the estimate dropped and could claim completeness. - An occurrence is now an attempt-ledger entry (workflow_run_id, source_event_sequence). A usage event belongs to the occurrence of its workflow run, Smithers task, iteration, and attempt whose start and terminal events span its sequence; each occurrence is priced from its own latest snapshot and an executed agent occurrence with none is imputed, across every workflow run of the run. The Smithers task ID is `node:<strategy_attempt_id>`, which every task manifest enforces. - `spend_estimate.unaccounted_attempts.entries` carry `workflow_run_id` and `source_event_sequence` and are unique by them (schema, semantic check, builder, docs). - Default-usage imputation prices catalog models at base rates; the 2,000,000-token default total no longer selects a context tier. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…v4 route lookup (#1268) Accounting v4 records pricing_catalog.status "available" when the catalog was fetched (or no model had a route) yet some models stay unresolved; route-exclusive pricing makes that common. The Modal worker-result gate rejected that state, failing status.json/result.json writes and reads. Drop the rule; uncosted events of an unresolved model still require usage.partial_pricing through the existing unpriced-event check, and the unavailable/disabled consistency rules stay. The #1268 changelog bullet now says v4 keeps its labels and semantics but looks up newly priced models through the route-exclusive catalog, so some models an aggregator used to price are now unresolved in v4. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…c path redaction Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| const occurrences = spendEstimateAttemptOccurrences(replayNodeAttempts(input.layout).entries, input.usageEntries); | ||
| const unaccountedAttempts = occurrences.unaccounted; | ||
| const previous = input.metadata.spend_estimate; | ||
| if (input.usageEntries.length === 0 && unaccountedAttempts.length === 0) return { changed: previous !== undefined }; |
There was a problem hiding this comment.
Source spend disappears When a continuation has a source run but no local usage or executed agent attempts, this return skips the source-run estimate. The continuation’s synchronized
run.json then has no spend_estimate or spend lineage, even if the source run has a persisted estimate.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/workflow-sync.ts
Line: 1493
Comment:
**Source spend disappears** When a continuation has a source run but no local usage or executed agent attempts, this return skips the source-run estimate. The continuation’s synchronized `run.json` then has no `spend_estimate` or spend lineage, even if the source run has a persisted estimate.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| entry: { row: unknown; path: string }, | ||
| source: Readonly<Record<string, unknown>> | ||
| ): SemanticGateIssue[] { | ||
| return source.recommendation === undefined && at(entry.row, ["recommendation"]) !== undefined |
There was a problem hiding this comment.
Rewritten remediation can publish If a source finding has a recommendation, this new error check does not catch a report row that changes it. The existing preservation check issues only a warning, so the report can publish different advice under
### Remediation despite the carry-only contract.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/artifacts/src/semantic-gates.ts
Line: 924
Comment:
**Rewritten remediation can publish** If a source finding has a recommendation, this new error check does not catch a report row that changes it. The existing preservation check issues only a warning, so the report can publish different advice under `### Remediation` despite the carry-only contract.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements #1268. The final report now names the evaluated commit, drops low-level diagnostics from
report.md, renders remediation for every production finding with readable inline code, and always shows a numericEstimated spend. The prompt, report contract and projection, renderer, gates, docs, and regression tests change together.Commit provenance. The
Source run IDrow is replaced by- Commit: <full hash>, read from the sealed data-governance record (target.commit).report.jsonrun_metadata.target_commitis required: a 40/64-hex string, ornullwhen no Git commit was recorded (rendered as- Commit: `none` (no Git commit was recorded for the evaluated target)). A missing or invalid governance record fails the report task (artifact-contract failure) instead of using a placeholder.source_run_id,source_run_ids) and the dirty and worktree digests stay structured-only.Simpler Markdown.
## Scoped coverage evidenceand## Artifact validation warningsare no longer rendered.report.json.coverage_evidenceandrun_metadata.artifact_validation_warningsare unchanged.artifact-validation-warnings.{json,md}companions are byte-identical. The public bundle now requires both companions whenever the report carries warnings.REPORT_COVERAGE_EVIDENCE_MARKDOWN_UNEXPECTED) instead of requiring it.Remediation and readable code.
### Remediation, after the Proof of Concept and any family variants. It shows the finding's carriedrecommendation, or this fixed sentence: "No remediation was recorded for this finding, and Ultrafuzz does not infer one. Confirm the root cause in the description and Proof of Concept before designing a fix."recommendation. A new semantic-gate error fires when a report row adds one that its source finding lacks. Producers are asked to set it only when their cited evidence establishes the fix, and dedupe keeps the root's.findingProseinstead ofpublicProse. Backtick spans render as inline code and intraword_stays raw. Link, definition, list, heading, HTML, and fence injection stay neutralized._and code spans.publicProseis unchanged for the coverage producer and blocker summaries, whose gates compare bytes.Always-numeric spend estimate. A new
run.json#spend_estimate(ultrafuzz.spend-estimate.v1) is rebuilt on every sync. It prices each attempt from, in order:ultrafuzz.fallback-pricing.2026-10-01, for missing rates or usage;Lineage comes from the source run's persisted estimate. A no-op sync never rewrites
run.json.openrouter/is stripped;vendor/modeland~ids are priced only from OpenRouter;claude-,gpt-/o*,deepseek, andkimiids are priced only from their first-party entries.[1m]-style aliases are stripped, and zero rates are ignored unless the id ends in:free.spend_estimate(or the live estimate) plus an imputed estimate of the report's own production. Tokens, spend, andpartial_pricingcome from one source. This fixes the old loss ofpartial_pricing: truewhen cumulative spend wasunavailable.spend_estimate, which includes report-production usage.Nothing shows
+,unavailable, orpartial_pricinginreport.md.run_metadata.estimated_spendmust match^\$(?:0|[1-9][0-9]*)\.[0-9]{2,10}$. Method, price sources (catalog provider and model id), fallback rates snapshotted per model, assumption codes, unaccounted attempts,complete, andupdated_atlive inspend_estimate.ultrafuzz reportdiagnostics compare runtime presentations to therun.jsonthey were built from and no longer expect+.Accounting v4 keeps its labels (
$x+andunavailable) and semantics forstats, evals, and Modal. Its catalog lookup for newly fetched models is now route-exclusive. The Modal worker-result gate now acceptsavailablepricing with unresolved models, a documented v4 state that the route change makes more common.Sample
report.md(rendered by this branch's renderer)Breaking changes
ultrafuzz/report@3and therun.jsonschema change in place, following the #1095, #1120, and #1167 precedent (see the CHANGELOG entry):target_commit, spend pattern).report.md, so scores across this boundary are not comparable..ultrafuzz/prompts/review/final-report.mdand rerunultrafuzz init.Design decisions worth reviewing
recommendationis carried, never authored, by the report stage. This keeps Remediation traceable to the producer that saw the code. The cost is that most reports will show the fallback sentence until producer lanes adopt the newfindings.mdxinstruction.partial_pricing: true, because it always contains an imputed estimate of its own production. The terminal presentation replaces that estimate with recorded usage.\,>,<, a word-starting|, or a run such as/*, so finding text that renders more literally cannot reveal one.Known limitations / follow-ups (not changed here)
<(e.g.cat </srv/x) already published unredacted at base, and still does. It needs its own redaction decision.ULTRAFUZZ_CACHE_READ_RATIOis still undocumented.Validation
All runs used Node 24.9, locally:
pnpm -w format:check,pnpm -w lint,pnpm -w build,pnpm -w knip,node scripts/smithers-patches.mjs --check,ESLINT_PLUGIN_DIFF_COMMIT=5c1a9775 pnpm -w lint:strict:ci,pnpm -w size,pnpm -w docs:check,pnpm -w typecheck.runtime.test.js(12 shards): 356/357.5c1a9775on this Node version: "a missing or unrenderable prompt fails only its own task at assert-task-inputs" and "generated Smithers verification marker root must be a canonical directory".scripts/ci79/79.Closes #1268.
🤖 Generated with Claude Code
The PR is not ready to merge until source-only continuations retain spend lineage and published remediation remains identical to its source recommendation.
Fix with agent prompt
Summary
The PR adds evaluated-commit provenance and a persisted spend estimate, simplifies report Markdown, and renders source-finding recommendations as remediation.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Usage and attempt ledgers] --> B[Spend estimate] S[Source run estimate] --> B B --> R[run.json] R --> P[Report-start projection] F[Source findings] --> G[Preservation gate] G --> M[Remediation in report.md] P --> MReviews (1) · Last reviewed commit: "docs: describe block-start escaping, tab..."