diff --git a/CHANGELOG.md b/CHANGELOG.md index f7ee7d34a..2370255d0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -29,6 +29,7 @@ ### Other changes +- **[runtime] [docs]** `ultrafuzz run --run-id ` now fails with `RUN_ALREADY_EXISTS` before planning when the workflow engine already records a run with that ID, as it does after `ultrafuzz clean ` (#1258). Before, `run` rebuilt the plan and the execution snapshot, about 6 minutes and 1 GB for Damn Vulnerable DeFi, then failed at submission with `WORKFLOW_SUBMISSION_FAILED` / `DETACHED_ADMISSION_FAILED` / `RUN_EXISTS`, and left a partial run directory behind. The check asks the engine only once the project root has the engine's database, and any answer other than "this run exists" lets the launch go ahead as before. `clean` still leaves the engine record behind (#1257), so the ID stays taken. - **[config] [runtime] [prompts] [docs]** Adds `run.friction_log_enabled` (default `false`). When enabled, agent tasks are told to record Ultrafuzz, tooling, or instruction roadblocks with [Frog](https://github.com/wevm/frog) (`frog@1.1.0`, now a pinned dependency of `@ultrafuzz/runtime`) through a run-local `/friction-bin/ultrafuzz-friction-log` command, which runs the Frog of the Ultrafuzz install that rendered the workflow. The command accepts only `log` and `list` with a few local options, refuses the built-in flags of incur, Frog's CLI framework, such as `--mcp`, where an option value belongs, points Frog at `/friction` with `GIT_DIR` set to a path that does not exist, so entries land in `/friction/.agents/friction-log/`, unsets Frog's GitHub and Postgres store variables and incur's `COMPLETE`, points `GH_CONFIG_DIR` at a path that does not exist, sets `NO_UPDATE_NOTIFIER=1`, and gives Frog `/dev/null` as stdin. It guards against misuse by mistake and enforces nothing: agents run unsandboxed (see `docs/security.md`), so `/friction` is only where the command writes, and it removes no GitHub credential, because Frog falls back to `gh auth token`, which still finds a token kept in the system keyring. Every task preparation creates the directory and rewrites the command, best-effort: a friction log that cannot be prepared never fails the task. Frog is installed with the runtime even when the friction log is disabled, but runs leave it out of their trusted CLI closure and execution snapshot. Its `postgres` dependency moves the install paths of the Smithers packages that reach drizzle-orm (see the `resume` change above); the engine still runs with `SMITHERS_BACKEND=sqlite`. A disabled run renders byte-identical prompts. Entries are free-form agent text that the artifact secret scan does not cover, so check them for credentials, such as RPC URLs with API keys, before publishing anything. Read them as Markdown, or with the Frog pinned in the Ultrafuzz checkout and `GIT_DIR` set to a path that does not exist (see `docs/reference/configuration.md`), never a bare `npx frog`, and delete a malformed entry's directory, since Frog refuses every `log` and `list` while one is malformed (#172). - **[runtime] [docs]** A failed `ClaudeAgent` or `DeepSeekAgent` attempt now records the failure Claude Code states in its result, such as a contended OAuth refresh or a rejected DeepSeek API key, in the node's `last_error` and the attempt ledger's `failure_message`, for example `Claude run failed See https://smithers.sh/reference/errors (agent stated: Failed to refresh OAuth token: …)`. Before, the record was only the generic `Claude run failed` message, and the cause survived only in the Claude Code session transcript. `ultrafuzz inspect --json` shows it as `data.state.nodes..last_error`, the dashboard shows it as the node's latest error, and a public Modal eval worker's `public-eval-diagnostics.json` carries it as the failed node's `failure_message`; `ultrafuzz status` and `ultrafuzz why` do not show it, because they relay the workflow engine's own summaries, which carry the error message but not its details. The thrown error Smithers classifies is unchanged, so quota parking, the auth disable and the retry behaviour from #1171 are unaffected; the statement travels as `details.agentStatedFailure` and gets the same secret redaction and length cap as `failure_message`. `ultrafuzz run` refuses a project whose stock `.smithers/agents` files predate this release (`CONTROLLER_SOURCE_UNTRUSTED`), so existing projects must re-run `ultrafuzz init`, which keeps an existing `ultrafuzz.toml`, topology and prompts, before their next run. A run launched by an earlier release records the statement only after `ultrafuzz resume --refresh-controller`, which continues it with this release's workflow and adapters (#1084). - **[modal]** Moves `@grpc/grpc-js` to 1.14.5, past the High advisory GHSA-m9gg-hp2v-232j (`getAuthContext` could return unauthorized certificates as authorized in certain configurations). diff --git a/docs/reference/cli.md b/docs/reference/cli.md index c8299c5da..9968648a8 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -231,6 +231,14 @@ replacement model. `--audit-profile` selects a profile for one command, while `--topology-path` atomically replaces the project or profile topology for that command. +A run ID must be new to the project. `run` fails with `RUN_ALREADY_EXISTS` +before it plans anything when the run directory exists, or when the workflow +engine already records a run with that ID. The second case happens after +`ultrafuzz clean `, which removes the run directory but not the engine's +record. Choose another `--run-id`, or continue a run whose directory still +exists with `ultrafuzz resume `. `run` asks the engine only once the +project has engine records, so a project's first launch skips the query. + ## Audit Profiles and Packaged Topologies ```bash diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index 46ec3ed67..77f234de8 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -3,6 +3,7 @@ import path from "node:path"; import { appendEvent, + isRecord, assertRunPlanDocument, assertPlannedGraph, assertPlannedGraphSemantics, @@ -72,6 +73,7 @@ import { assertSmithersControllerRefreshable, renderCurrentSmithersController, requestSmithersPause, + runSmithersInspectionCommand, runSmithersLifecycleCommand, type SmithersResumeInspection, assertSealedDataGovernance, @@ -189,6 +191,8 @@ function controllerRefreshInspectionEnvironment( } export async function startRun(input: StartRunInput) { + const knownRun = await workflowRunAlreadyRecorded(input); + if (knownRun !== undefined) return runtimeFailure([knownRun]); let createdLayout: RunLayout | undefined; const planned = await planRun(input, { enforceDataGovernance: true, @@ -722,6 +726,46 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { } } +/** + * Refuse a run ID the workflow engine already has a run for, before planning builds the run's plan and + * execution snapshot (#1258). The engine rejects such a launch only at submission, with `RUN_EXISTS`, + * after minutes of preparation and with a partial run directory left behind; `clean` removes the run + * directory but not the engine's record. The engine creates its database in the project root on the + * first launch, so a project without one has no runs to collide with and is not asked. Any answer + * other than the engine reporting this exact run lets the launch go ahead, as before: submission + * still rejects a duplicate. + */ +async function workflowRunAlreadyRecorded(input: StartRunInput): Promise { + if (input.runId === undefined) return undefined; + let runId: string; + try { + runId = validateSafeId(input.runId, "run ID"); + } catch { + // Planning reports the invalid ID. + return undefined; + } + const projectRoot = path.resolve(input.projectRoot); + if (!fs.existsSync(path.join(projectRoot, "smithers.db"))) return undefined; + const workflowRunId = `ultrafuzz-${runId}`; + const inspected = await runSmithersInspectionCommand({ + args: ["inspect", workflowRunId, "--format", "json"], + projectRoot, + env: input.env + }); + // The runner prints a found run's inspection bare, and the full-output envelope wraps it in `data`. + const report = inspected.json; + const data = isRecord(report) && report.ok === true ? report.data : report; + const run = inspected.ok && isRecord(data) ? data.run : undefined; + if (!isRecord(run) || run.id !== workflowRunId) return undefined; + const status = typeof run.status === "string" ? ` (${run.status})` : ""; + return { + code: "RUN_ALREADY_EXISTS", + message: `run ${runId} already exists in the workflow engine's records${status}, so it cannot be launched again; choose another run ID, or continue the existing run with \`ultrafuzz resume ${runId}\` if its run directory still exists`, + severity: "error", + source: "runtime" + }; +} + // A missing or unreadable config fails the resume: without it the run's execution mode is unknown. function readContinuationResolvedConfig(runRoot: string, configPath: string, runId: string): ResolvedConfig { try { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index f380f27d7..a859f5679 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -27206,6 +27206,81 @@ test("a pending lifecycle link reconciles split source and target projections fr assert.equal(reconciledJournal.entries?.at(-1)?.phase, "committed"); }); +// #1258: a run ID the engine already has a run for is refused before planning builds the plan and the +// execution snapshot, instead of at submission minutes later with a partial run directory left behind. +// The pinned runner prints a found run's `inspect --format json` bare; the full-output envelope wraps it. +for (const shape of ["bare", "enveloped"] as const) { + test(`run refuses a run ID the workflow engine already records before planning anything (${shape} inspection)`, async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = `engine-known-run-${shape}`; + const inspection = workflowInspect({ + workflowRunId: `ultrafuzz-${runId}`, + status: "finished", + state: "succeeded", + steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] + }) as { data: unknown }; + const env = fakeLifecycleSmithersEnv(project, { inspect: shape === "bare" ? inspection.data : inspection }); + // The engine keeps its run records here; `clean` removed the run directory but not the record. + fs.writeFileSync(path.join(project, "smithers.db"), ""); + + const refused = await startRun({ projectRoot: project, runId, env }); + + assert.equal(refused.ok, false); + const diagnostic = refused.diagnostics.find(({ code }) => code === "RUN_ALREADY_EXISTS"); + assert.ok(diagnostic, JSON.stringify(refused.diagnostics)); + assert.match(diagnostic.message, /already exists in the workflow engine's records \(finished\)/u); + assert.doesNotMatch(diagnostic.message, /smithers/iu); + assert.equal(fs.existsSync(path.join(project, ".ultrafuzz", "runs", runId)), false); + const commands = fs.readFileSync(env.SMITHERS_FAKE_LOG ?? "", "utf8"); + assert.match(commands, new RegExp(`^inspect ultrafuzz-${runId} --format json$`, "mu")); + assert.doesNotMatch(commands, /^up /mu); + }); +} + +test("run asks the engine about its run ID only once the engine has records, and launches a new one", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "engine-new-run"; + const env = fakeLifecycleSmithersEnv(project, { + inspect: { + ok: false, + error: { code: "RUN_NOT_FOUND", message: `Run not found: ultrafuzz-${runId}` }, + meta: { command: "inspect", duration: "1ms" } + } + }); + const commandLog = env.SMITHERS_FAKE_LOG; + assert.ok(commandLog); + fs.writeFileSync(path.join(project, "smithers.db"), ""); + + const launched = await startRun({ projectRoot: project, runId, env }); + + assert.equal( + launched.diagnostics.some(({ code }) => code === "RUN_ALREADY_EXISTS"), + false, + JSON.stringify(launched.diagnostics) + ); + const commands = fs.readFileSync(commandLog, "utf8"); + assert.match(commands, new RegExp(`^inspect ultrafuzz-${runId} --format json$`, "mu")); + assert.match(commands, /^up /mu); + + // Without engine records there is nothing to collide with, and the engine is not asked. + const fresh = tempProject(); + initProject({ projectRoot: fresh, force: true }); + writeSmallTopology(fresh); + const freshEnv = fakeLifecycleSmithersEnv(fresh, { + inspect: workflowInspect({ workflowRunId: "ultrafuzz-fresh-run", status: "running", steps: [] }) + }); + const started = await startRun({ projectRoot: fresh, runId: "fresh-run", env: freshEnv }); + assert.equal(started.ok, true, JSON.stringify(started.diagnostics)); + assert.doesNotMatch( + fs.readFileSync(freshEnv.SMITHERS_FAKE_LOG ?? "", "utf8"), + /^inspect ultrafuzz-fresh-run --format json$/mu + ); +}); + test("ordinary resume checks active-run ownership before detached preflight", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true });