Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@
- **[runtime] [docs]** The `OpenRouterAgent` provider configuration now sets Codex's `stream_max_retries = 10` and `stream_idle_timeout_ms = 1800000`, up from Codex's defaults of 5 and 300000 (#676). A dropped or idle response stream failed the node with `stream disconnected before completion`, and the node retry started the session over. Codex reconnects by re-sending the current request in the same session, so a reconnect keeps the agent's earlier work. Codex's idle timer counts only stream events, and OpenRouter's keepalive comments do not reset it, so a model that streams no reasoning events could stay silent past the 5-minute default and drop on every reconnect. A genuine outage still fails within about 3.5 minutes. Provider route IDs do not change. `CodexAgent` on Codex's built-in OpenAI provider is unchanged, because Codex rejects overrides of built-in provider IDs.
- **[runtime] [docs]** `ultrafuzz run --run-id <id>` now fails with `RUN_ALREADY_EXISTS` before planning when the workflow engine already records a run with that ID, as it does after `ultrafuzz clean <id>` (#1258). Before, `run` rebuilt the plan and the execution snapshot, about 6 minutes and 1 GB for Damn Vulnerable DeFi, then failed at submission with `WORKFLOW_SUBMISSION_FAILED` / `DETACHED_ADMISSION_FAILED` / `RUN_EXISTS`, and left a partial run directory behind. The check asks the engine only once the project root has the engine's database, and any answer other than "this run exists" lets the launch go ahead as before. `clean` still leaves the engine record behind (#1257), so the ID stays taken.
- **[runtime] [docs]** `validate`, `plan` and `run` now report a `TOPOLOGY_TIMEOUT_SHADOWS_DEFAULT` warning for each topology `timeout_seconds` pin, on a node or in a group's defaults, that is below the model profile `timeout_seconds` or `run.default_timeout_seconds` it overrides (#675). A pin wins even when it is the shorter window, so raising a default silently did not reach a pinned node: the packaged `goals`, `strategies`, `specialists` and `review` groups pin 7,200 seconds, and goal nodes kept timing out at two hours after a longer default was configured. The warning names the pin, the default it overrides and the nodes it applies to. Before, only a CI test on the packaged topologies checked this, so a project's own topology got no check.
- **[runtime] [modal] [docs]** `resume --retry-failed` no longer reruns a failed task in a `failure_policy: continue` group, such as a property lens or an optional strategy, once a task that treats it as optional has started without it (#1231). The rerun could not reach the outputs that task had already produced, yet once it succeeded every planned task read as succeeded, so the final report said COMPLETE and a `--require-complete` run could end `succeeded`. The task now stays failed, the report stays PARTIAL, `resume` reports a `WORKFLOW_RETRY_SKIPPED` warning that names it and the tasks that ran without it, and other failed tasks are retried as before. The Modal benchmark worker no longer resumes a finished run whose only failures are such tasks, which it did over every failed task and then waited for a change that would not come.
- **[runtime] [modal] [docs]** `resume --retry-failed` no longer reruns a failed task in a `failure_policy: continue` group, such as a property lens or an optional strategy, once a task that treats it as optional has started without it (#1231). The rerun could not reach the outputs that task had already produced, yet once it succeeded every planned task read as succeeded, so the final report said COMPLETE and a `--require-complete` run could end `succeeded`. The task now stays failed, the report stays PARTIAL, `resume` reports a `WORKFLOW_RETRY_SKIPPED` warning that names it and the tasks that ran without it, and other failed tasks are retried as before. A task whose preparation failed has not started: its preparation, which admits its inputs, reruns after the failed task, so both are retried. A task whose agent or verifier failed reruns at once, without waiting for the failed task, so it counts as started. The Modal benchmark worker no longer resumes a finished run whose only failures are such tasks and the tasks skipped because one of them failed, such as the later stateful-invariant stages after a failed `stateful-invariant-setup`, which it did over every failed or skipped task and then waited for a change that would not come.
- **[config] [runtime] [prompts] [docs]** Adds `run.friction_log_enabled` (default `false`). When enabled, agent tasks are told to record Ultrafuzz, tooling, or instruction roadblocks with [Frog](https://github.com/wevm/frog) (`frog@1.1.0`, now a pinned dependency of `@ultrafuzz/runtime`) through a run-local `<run>/friction-bin/ultrafuzz-friction-log` command, which runs the Frog of the Ultrafuzz install that rendered the workflow. The command accepts only `log` and `list` with a few local options, refuses the built-in flags of incur, Frog's CLI framework, such as `--mcp`, where an option value belongs, points Frog at `<run>/friction` with `GIT_DIR` set to a path that does not exist, so entries land in `<run>/friction/.agents/friction-log/`, unsets Frog's GitHub and Postgres store variables and incur's `COMPLETE`, points `GH_CONFIG_DIR` at a path that does not exist, sets `NO_UPDATE_NOTIFIER=1`, and gives Frog `/dev/null` as stdin. It guards against misuse by mistake and enforces nothing: agents run unsandboxed (see `docs/security.md`), so `<run>/friction` is only where the command writes, and it removes no GitHub credential, because Frog falls back to `gh auth token`, which still finds a token kept in the system keyring. Every task preparation creates the directory and rewrites the command, best-effort: a friction log that cannot be prepared never fails the task. Frog is installed with the runtime even when the friction log is disabled, but runs leave it out of their trusted CLI closure and execution snapshot. Its `postgres` dependency moves the install paths of the Smithers packages that reach drizzle-orm (see the `resume` change above); the engine still runs with `SMITHERS_BACKEND=sqlite`. A disabled run renders byte-identical prompts. Entries are free-form agent text that the artifact secret scan does not cover, so check them for credentials, such as RPC URLs with API keys, before publishing anything. Read them as Markdown, or with the Frog pinned in the Ultrafuzz checkout and `GIT_DIR` set to a path that does not exist (see `docs/reference/configuration.md`), never a bare `npx frog`, and delete a malformed entry's directory, since Frog refuses every `log` and `list` while one is malformed (#172).
- **[runtime] [docs]** A failed `ClaudeAgent` or `DeepSeekAgent` attempt now records the failure Claude Code states in its result, such as a contended OAuth refresh or a rejected DeepSeek API key, in the node's `last_error` and the attempt ledger's `failure_message`, for example `Claude run failed See https://smithers.sh/reference/errors (agent stated: Failed to refresh OAuth token: …)`. Before, the record was only the generic `Claude run failed` message, and the cause survived only in the Claude Code session transcript. `ultrafuzz inspect <run-id> --json` shows it as `data.state.nodes.<node-id>.last_error`, the dashboard shows it as the node's latest error, and a public Modal eval worker's `public-eval-diagnostics.json` carries it as the failed node's `failure_message`; `ultrafuzz status` and `ultrafuzz why` do not show it, because they relay the workflow engine's own summaries, which carry the error message but not its details. The thrown error Smithers classifies is unchanged, so quota parking, the auth disable and the retry behaviour from #1171 are unaffected; the statement travels as `details.agentStatedFailure` and gets the same secret redaction and length cap as `failure_message`. `ultrafuzz run` refuses a project whose stock `.smithers/agents` files predate this release (`CONTROLLER_SOURCE_UNTRUSTED`), so existing projects must re-run `ultrafuzz init`, which keeps an existing `ultrafuzz.toml`, topology and prompts, before their next run. A run launched by an earlier release records the statement only after `ultrafuzz resume <run-id> --refresh-controller`, which continues it with this release's workflow and adapters (#1084).
- **[modal]** Moves `@grpc/grpc-js` to 1.14.5, past the High advisory GHSA-m9gg-hp2v-232j (`getAuthContext` could return unauthorized certificates as authorized in certain configurations).
Expand Down
8 changes: 6 additions & 2 deletions docs/reference/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -373,8 +373,12 @@ task was admitted while the failed one had no output, and a rerun cannot reach
the outputs it already produced, so the run would read as complete over work
that never used the rerun's output. `resume` leaves such a task failed, reports
a `WORKFLOW_RETRY_SKIPPED` warning that names it and the tasks that ran without
it, and the final report stays partial. The Modal benchmark worker does not
resume a finished run whose only failures are such tasks.
it, and the final report stays partial. A task whose preparation failed has not
started: its preparation, which admits its inputs, reruns after the failed task,
so both are retried. A task whose agent or verifier failed reruns at once,
without waiting for the failed task, so it counts as started. The Modal
benchmark worker does not resume a finished run whose only failures are such
tasks and the tasks skipped because one of them failed.

`resume` runs the workflow engine from Ultrafuzz's own install: pnpm applies
the committed compatibility patches (`patches/`) to it at install time, so
Expand Down
7 changes: 6 additions & 1 deletion packages/modal/src/resume.ts
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,8 @@ export function modalDurableResumeCommand(cliPath: string, runId: string, projec
* `retainedFailures` names the failed tasks `resume --retry-failed` leaves failed because a started
* consumer already ran without them (#1231). A terminal run whose only failures are those has
* nothing a resume would rerun, and `waitForTerminalRun` would wait for a change that never comes.
* A skipped task is not one either: `--retry-failed` reruns it only by rerunning the failed task it
* waited on, which has a failed entry of its own.
*/
export function modalDurableRunNeedsResume(
state: ModalResumeRunState,
Expand All @@ -72,7 +74,10 @@ export function modalDurableRunNeedsResume(
if (counts.remaining > 0) return true;
if (counts.failed === 0) return false;
const failedNodeIds = Object.entries(state.nodes ?? {})
.filter(([, node]) => node.status !== undefined && CHECKPOINT_FAILED_STATUSES.has(node.status))
.filter(
([, node]) =>
node.status !== undefined && node.status !== "skipped" && CHECKPOINT_FAILED_STATUSES.has(node.status)
)
.map(([nodeId]) => nodeId);
return failedNodeIds.length === 0 || failedNodeIds.some((nodeId) => !retainedFailures.has(nodeId));
}
Expand Down
22 changes: 22 additions & 0 deletions packages/modal/test/resume.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -291,6 +291,28 @@ describe("Modal durable evaluation resume", () => {
).toBe(true);
});

it("does not resume a terminal run whose other failures are skipped behind a retained one", () => {
// The review ran without the failed setup, so `--retry-failed` leaves it failed, and the later
// stages of its group, which require it, stay skipped.
const state = {
run_id: "durable-run-one",
status: "succeeded",
nodes: {
"stateful-invariant-setup": { status: "failed" },
"stateful-invariant-handlers": { status: "skipped" },
"stateful-invariant-campaign": { status: "skipped" },
"dedupe-findings": { status: "succeeded" }
}
};
const counts = { succeeded: 1, failed: 3, remaining: 0 };
const retained = new Set(["stateful-invariant-setup"]);
expect(modalDurableRunNeedsResume(state, counts, undefined, retained)).toBe(false);
// Rerunning the setup reruns the stages skipped behind it.
expect(modalDurableRunNeedsResume(state, counts)).toBe(true);
const otherFailure = { ...state, nodes: { ...state.nodes, "boundary-tests": { status: "failed" } } };
expect(modalDurableRunNeedsResume(otherFailure, { ...counts, failed: 4 }, undefined, retained)).toBe(true);
});

it("resumes operational checkpoints but never retries a terminal task outcome", () => {
expect(
modalDurableRunNeedsResume(
Expand Down
19 changes: 16 additions & 3 deletions packages/runtime/src/retry-failed-omissions.ts
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,17 @@ export function failedProducersStartedConsumersOmitted(
}

const SMITHERS_FAILED_STATES = new Set<SmithersNodeState>(["failed", "stalled"]);
const SMITHERS_STARTED_STATES = new Set<SmithersNodeState>(["in-progress", "finished", "failed", "stalled"]);
const SMITHERS_ADMITTED_STATES = new Set<SmithersNodeState>(["in-progress", "finished"]);

/** {@link failedProducersStartedConsumersOmitted}, judged from the Smithers node states a resume inspects. */
/**
* {@link failedProducersStartedConsumersOmitted}, judged from the Smithers node states a resume inspects.
*
* A consumer starts with its preparation, which waits for its producers and admits its optional
* inputs. `--retry-failed` reruns a failed or stalled preparation behind the producer's rerun, so
* that consumer can still read it. A failed agent or verifier reruns at once on its finished
* preparation, without waiting for the producer. An interrupted preparation still counts, which
* errs toward leaving the producer failed.
*/
export function failedProducersStartedConsumersOmittedInWorkflow(
tasks: readonly SmithersTaskManifestTask[],
nodeStates: ReadonlyMap<string, SmithersNodeState>
Expand All @@ -56,11 +64,16 @@ export function failedProducersStartedConsumersOmittedInWorkflow(
return failedProducersStartedConsumersOmitted(
tasks,
(task) => nodesInState(task, SMITHERS_FAILED_STATES),
(task) => nodesInState(task, SMITHERS_STARTED_STATES)
(task) => {
const preparation = nodeStates.get(task.preparationSmithersNodeId);
return preparation !== undefined && SMITHERS_ADMITTED_STATES.has(preparation);
}
);
}

const RUN_FAILED_STATUSES = new Set<NodeStatus>(["failed", "timed-out"]);
// Run state cannot tell a failed preparation from a failed agent, so a failed consumer counts as
// started. Naming more failures than `resume` leaves can only skip a resume, never start an empty one.
const RUN_STARTED_STATUSES = new Set<NodeStatus>(["running", "succeeded", "failed", "timed-out"]);
const RUN_SUCCEEDED_STATUSES = new Set<NodeStatus>(["succeeded", "reused-from-prior-run"]);

Expand Down
60 changes: 60 additions & 0 deletions packages/runtime/test/runtime.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -28003,6 +28003,66 @@ for (const fixture of [
}
}

// A consumer admits its optional inputs in its preparation, which waits for its producers. A failed
// preparation reruns behind the producer's rerun and can read it, so the producer is retried too. A
// failed agent or verifier reruns at once on the preparation that already ran without the producer,
// so the producer stays failed (#1231).
for (const fixture of [
{ failed: "prepare:catalog", state: "failed", reset: "prepare:catalog", producerRetried: true },
{ failed: "prepare:catalog", state: "stalled", reset: "prepare:catalog", producerRetried: true },
{ failed: "node:catalog", state: "failed", reset: "node:catalog", producerRetried: false },
{ failed: "verify:catalog", state: "failed", reset: "node:catalog", producerRetried: false }
] as const) {
test(`resume --retry-failed ${fixture.producerRetried ? "retries" : "leaves"} a failed continuing lens when its consumer's ${fixture.failed} is ${fixture.state}`, async () => {
const project = writeRetriedLensReviewProject();
const runId = `retry-omitted-lens-${fixture.failed.replace(":", "-")}-${fixture.state}`;
const preparationFailed = fixture.failed === "prepare:catalog";
const env = fakeLifecycleSmithersEnv(project, {
inspect: workflowInspect({
workflowRunId: `ultrafuzz-${runId}`,
status: "failed",
state: "failed",
steps: [
{ id: "node:lens", state: "failed", attempt: 1 },
{ id: "prepare:catalog", state: preparationFailed ? fixture.state : "finished", attempt: 1 },
...(preparationFailed
? [
{ id: "node:catalog", state: "skipped" as const, attempt: 0 },
{ id: "verify:catalog", state: "skipped" as const, attempt: 0 }
]
: fixture.failed === "node:catalog"
? [{ id: "node:catalog", state: fixture.state, attempt: 1 }]
: [
{ id: "node:catalog", state: "finished" as const, attempt: 1 },
{ id: "verify:catalog", state: fixture.state, attempt: 1 }
])
]
})
});
const run = await startRun({ projectRoot: project, runId, env });
assert.equal(run.ok, true, JSON.stringify(run.diagnostics));
const commandLog = env.SMITHERS_FAKE_LOG;
assert.ok(commandLog);
fs.writeFileSync(commandLog, "", "utf8");

const resumed = await resumeRun({ projectRoot: project, runId, force: true, retryFailed: true, env });

assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics));
const commands = fs.readFileSync(commandLog, "utf8");
assert.match(commands, new RegExp(`^timetravel .* --node-id ${fixture.reset} `, "mu"));
const producerReset = /^timetravel .* --node-id node:lens /mu;
const skipped = resumed.diagnostics.find((diagnostic) => diagnostic.code === "WORKFLOW_RETRY_SKIPPED");
if (fixture.producerRetried) {
assert.match(commands, producerReset);
assert.equal(skipped, undefined, JSON.stringify(resumed.diagnostics));
} else {
assert.doesNotMatch(commands, producerReset);
assert.ok(skipped, JSON.stringify(resumed.diagnostics));
assert.match(skipped.message, /lens: catalog already ran without it/u);
}
});
}

// `--reset-node` reruns a failed continuing task even when `--retry-failed` would leave it, so the
// resume does not also report that task as left failed.
test("resume --retry-failed --reset-node does not report the reset task as skipped", async () => {
Expand Down
Loading