Repository navigation
[rushd][Epic] rushd — server-first re-architecture (daemon + thin clients) #5894
Description
Activity
- added sub-issues
on Jul 22, 2026 - changed the title
[-][rushd][Epic] Rush 6 / rushd — server-first re-architecture (daemon + thin clients)[/-][+][rushd][Epic] rushd — server-first re-architecture (daemon + thin clients)[/+]on Sep 4, 2026 WS3/WS4 implementation handoff - September 8, 2026
The implementation for #5898 and #5899 is delivered in #6018 at
ba5753a9f146325deccb49e5aea485c6ba5aa3e2. Hosted run34217211265passed all six configurations: Ubuntu Node 20/22/24/26 and Windows Node 24/26. Every job completed package-manager integration and both second-checkout/build-cache passes. CodeQL and license/CLA are also green on this exact head.This includes native execution and safe lifecycle/restart handling, opt-in standalone clients and graph controls, warm-set ownership/status, Windows native mutex and path parity, Linux descendant completion, and the coupled native purge correction. Existing
rush/rushxdefaults remain unchanged.The Reporter DAG is published/restacked, and all 13 remaining published heads currently have successful checks. Required independent approvals and dependency-ordered landing remain outstanding; #5987 is the next stack edge and #5990 the separate privacy root. #6018 remains draft to respect that order. R9 (#5982), R11 (#5984), and WS5 (#5900) retain their coordinated compatibility/release gates.
No protection bypass, unreviewed merge, default flip, or full-epic closure was performed. Implementation and acceptance work is complete; the remaining gates are maintainer/release decisions.
- added a commit that references this issue
on Sep 22, 2026 Plugin work items before Rush 6 (from dogfooding
rush-clientin a large monorepo)We're dogfooding the opt-in client (
rush-client, built with the snapshot workflow from #6046) in a large internal monorepo. It configures 14 autoinstaller plugins, 4 of them from this repo. Out of the box, everybuild/rebuildfalls back to in-process Rush because of the plugin gate:- On
main,PhasedCommandEngine.parseAsyncrejects any workspace that has external plugins configured (PhasedCommandEngine.ts:90). - [rush] Dogfood the Rush daemon in rushstack: allow command-scoped plugins, add snapshot workflow and guide #6046 narrows this to plugins that participate in the command. That still rejects 9 of the 14, because a plugin without
associatedCommandsalways counts as participating, even when it only tapsafterInstall.
The repo is fixing its own plugins separately. Below are the upstream items we think need to land on 5.x, before the Rush 6 cutover (#5900), for a repo like this to run the client with its plugins enabled. Line numbers are at 982b33b.
Plugin host (rush-lib / rush-daemon)
- Land [rush] Dogfood the Rush daemon in rushstack: allow command-scoped plugins, add snapshot workflow and guide #6046, so that plugins scoped only to other commands count as inert and are allowed.
- Add an explicit daemon-compatibility declaration for plugins that do participate. Today there is no opt-in. We're prototyping three sources: the plugin manifest
daemonCompatible: true(for the plugin author),rush.jsondaemon.compatiblePlugins(for the repo) andRUSH_DAEMON_COMPATIBLE_PLUGINS(env).- A declaration clears every gate reason, but only for the declared plugin.
- The declared set is part of daemon identity.
- A declared plugin's parameters stay in the engine parameter identity.
- This should converge with [rush reporter][R9] Complete Rush 6 reporter and plugin migration #5982's "gate incompatible plugins before
apply()", so that there is one compatibility mechanism rather than two.
- Ship the schema additions in a 5.x release first.
rush-plugin-manifest.schema.json(plugins[]) andrush.schema.jsonboth useadditionalProperties: false, and 5.178.x has nodaemonsection at all.- So a repo can't check in a declaration until it's on a release that accepts the keys, and a plugin that adds the manifest key fails to load on older Rush.
- Order: schema release, then repos bump
rushVersion, then plugins publish declarations. Lockstep plugins can ship the key in the same release as the schema.
- Make plugin code part of the daemon's runtime identity.
- Plugins load with
require()(PluginLoaderBase.ts:145), and the runtime fingerprint covers only Rush's own package (WorkspaceRequestLifecycle.ts:115). - After a pull that changes only plugin code, the daemon keeps running the old code. A tier-1 reload doesn't help, because the module is still in the require cache.
- Fix: add each loaded plugin's package folder (the realpath, so
link:targets count) as a restart-tier input. Test it by changing only a plugin's entry file between requests.
- Plugins load with
- Stop plugin writes to
process.envfrom breaking daemon identity.- The resolver compares the live
process.envwith its startup snapshot on every request (ProductionDaemonRequestResolver.ts:152). A plugin that setsprocess.env.Xin a session or iteration hook therefore makes every later requestunsupported. - Fix, either: treat keys written during plugin hooks as plugin-owned (excluded from the live check, still passed to operations); or detect the write and name the plugin in the diagnostic.
- The resolver compares the live
- Add per-request telemetry in engine mode.
- The engine-mode parser returns before it creates
Telemetry(RushCommandLineParser.ts:419, whileTelemetryis created at:511). - So daemon-served builds write no
common/temp/telemetryfile and never fireflushTelemetry, and plugins that report through that hook see nothing. - Needed: one record and one
flushTelemetryper served request, including coalesced participants and fully up-to-date requests. Timings should be relative to the request, not the daemon'sprocess.uptime().
- The engine-mode parser returns before it creates
- Add a request-scoped hook.
- A fully up-to-date request schedules no iteration (
PhasedRequestRouter.ts:460), sobeforeExecuteIterationAsync,afterExecuteIterationAsyncandbeforeLognever fire. - As a result, plugins that produce per-command output (summaries, telemetry) silently skip exactly the warm requests.
- Fix: add a hook that carries the request's selection, results and timings, or at least document this behavior.
- A fully up-to-date request schedules no iteration (
- Add async dispose for plugins that hold resources.
PhasedCommandEngine's remarks already call for "request-scoped initialization and asynchronous disposal contracts". For example,rush-serve-pluginstarts HTTP/WebSocket servers for the commands it is configured for. This is needed before the daemon serves watch-style commands with plugins. - Document the daemon plugin contract in the plugin migration guide (WS5, [rushd][WS5] Rush 6 cutover: testing, observability & migration #5900):
applyandrunPhasedCommandrun once per engine.- Graph hooks run once per iteration, and one iteration can serve several coalesced requests.
- Up-to-date requests get no iteration.
- Plugins must not mutate
process.env,process.exitCodeor global timers. We found a plugin that wraps the globalsetTimeoutand later clears every pending timer; inside the daemon, that would also clear the daemon's idle and queue timers. - No interactive prompts inside the daemon. Fail fast with a message that is routed to the requesting client.
- Use the request's environment.
- Dispose through
abortControllerandcloseRunnersAsync. - How to declare compatibility.
- Diagnostics:
rush-client daemon status(or the fallback note) should list each plugin as inert, declared or rejected, with the reason. Then repo owners can see which plugin keeps them on the in-process path.
Plugins in this repo (audit each against the contract, then declare it in its manifest once the schema ships)
-
rush-resolver-cache-plugin: taps onlyafterInstall, so it is inert for phased commands. -
rush-buildxl-graph-plugin: taps onlyrunPhasedCommand.for(<buildXLCommandNames>). -
rush-serve-plugin: taps onlyrunPhasedCommand.for(<phasedCommands>). It needs the dispose contract before the daemon serves watch commands. -
rush-bridge-cache-plugin: tapsrunAnyPhasedCommand, but acts only when its action parameter is passed. Confirm that this parameter is part of the engine parameter identity. - Cloud cache providers (
rush-azure-storage-build-cache-pluginwith its built-inrush-azure-interactive-auth-plugin,rush-amazon-s3-build-cache-plugin,rush-http-build-cache-plugin):- [rush] Dogfood the Rush daemon in rushstack: allow command-scoped plugins, add snapshot workflow and guide #6046's snapshot doesn't register them. They are
publishOnlyDependenciesin source, andPluginManagerregisters built-in plugins only fromdependencies, soazure-blob-storagefails with "Unexpected cache provider". - Under a long-lived daemon, credentials have to be refreshed per request, and no interactive auth may run inside the daemon.
- Our dogfooding hasn't exercised either of these yet.
- [rush] Dogfood the Rush daemon in rushstack: allow command-scoped plugins, add snapshot workflow and guide #6046's snapshot doesn't register them. They are
Related (not plugin-specific)
- The snapshot's
_RUSH_LIB_PATHis a source-layout realpath with no enclosingnode_modules/@microsoft/rush-lib. Bundled plugins that resolve@microsoft/rush-libby name from it fail withMODULE_NOT_FOUND. Published installs are unaffected. - Custom phased commands (for example, a repo's
test) aren't served yet; onlybuild/rebuildare. So the most common agent command still runs in-process.
- On
Daemon and client work items from the same dogfooding run
This follows my plugin list above. These items came up once the plugin gate was bypassed and the daemon served builds in the same ~1.8k-project monorepo, and in rushstack clones. Line numbers are at 982b33b. "Confirmed" means a second person reproduced the item independently. The rest have one detailed reproduction so far.
Scale and latency
- Warm-set maintenance is O(N²) and blocks the event loop (confirmed).
#evictIdleAsynccallsgetStatus()once per project (WorkspaceWarmSet.ts:357-360), andgetStatus()re-ranks every project each time (:177). With about 1.8k retained projects, one pass takes about 4.3 s of main-thread CPU. It runs after every request, and again every ~35 s while the daemon is idle. Requests, status calls and cancels that arrive during a pass wait for it. The retained set is the whole workspace, not just the closure the daemon served. - A hot tier-0 request is slower than in-process Rush at this scale (confirmed). For a 1-project no-op, the daemon takes about 5.3 s against about 2.3 s for in-process Rush.
rush-project.jsonis loaded twice per request without a cache, about 1 s each time. Rigs resolve with preserved symlinks, so every project gets its own profile path, and heft-config-file's path-keyed cache never hits.- The per-request inputs snapshot, which includes a full-repo
git status, takes about 2.5 s.
- Any content change under
common/config/**forces a full graph reload (confirmed). That includes files written by a build. A single project that lives undercommon/configtherefore makes every request reload. - The engine is pinned to one command and parameter set (confirmed). Switching between
buildandrebuild, or passing a custom flag once, recreates the engine and drops every warm result. - Head-of-line blocking (confirmed). A request that arrives while another request's iteration is running waits for that whole iteration, even when the two selections don't overlap.
Correctness
- A hot
buildruns the:incrementalscript and still writes the build cache (confirmed). The worktree and the cache entries end up with stale outputs. - Undetected in-place changes to a dependency's outputs poison the cache (confirmed). The consumer builds from the changed outputs, and its result is written to the build cache.
- A stale
error.loggets packed into cache entries.OperationMetadataManager.saveAsync(:90-130) ignores ENOENT when the source log is missing, so the copy from an older run stays in.rush/temp/operation/<phase>/. A later restore (:178) then recreates an.error.logfor a clean build. In-process Rush has the same bug; the daemon just restores more often, for example after warm-set eviction. Fix: delete the destination on ENOENT. - Operations aren't reaped after an unclean daemon exit (SIGKILL or OOM). Reclaim signals the dead daemon's process group, but each operation is spawned detached into its own session. The next daemon re-runs the operation while the orphan is still writing to the same output folder.
Client, admission and cancellation
- A burst of clients with no daemon running times out on startup.
StartupLockusesLockFile.tryAcquire. On every attempt, that runsps -p <pid> -o lstartfor the waiter itself and for every other waiter's lock file (LockFile.ts:71-89,:433,:479). That is O(N²)psspawns, and each one scans all of/proc.- Result: with 32 simultaneous clients, 19–22 of them hit the fixed 15 s startup deadline (
connectOrStartDaemon.ts:84) in 3 of 5 rounds. Each fell back to in-process Rush, and all but one per round then failed on the repo lock. - The one fallback that did get the lock made two requests that were already queued in the daemon fail with
routingFailed. - In a micro-benchmark, reading
/proc/<pid>/statfirst on Linux cut the mean lock attempt from 1.3 s to 1.3 ms at N=8, and from 13.9 s to 0.5 ms at N=32. - Suggested fixes: get the start time without
ps; while a live startup reservation exists, have waiters only poll the connection; and don't fall back to in-process Rush while a live daemon owns the workspace.
- False admission timeout. A build waiting behind a running build can fail with "not admitted within 30000ms". This happens when a graph input changes during the wait: the retry reuses the default deadline, which has already expired.
- Cancellation isolation. When one of two coalesced clients cancels, the operations that only the cancelled client needed keep running, and the other client's result waits for them.
- A zero-size pty (columns = 0) fails every daemon request (confirmed). The failure is at
RequestEnvelopeValidation.ts:33, with no fallback, while in-process Rush works in the same pty. Treat 0 as unknown. - A
rushVersionthe client doesn't bundle repeats the version lookup on every command (confirmed). Every command ends in the same fallback and costs about 3 s extra, or about 137 s when the registry is unreachable. Fix: cache the incompatible result, and put a hard deadline on the lookup. - Global flags before the command (
rush-client --quiet build) silently route to in-process Rush (confirmed). - A daemon shutdown during a build is reported as a user cancellation (exit 130, reason dropped) (confirmed).
- A served
rushxscript holds the workspace lifecycle gate for as long as it runs (confirmed). Builds that need a reload wait behind a dev server.
Environment and identity
- Request-scoped environment variables need their own channel. Today, session-scoped variables (IDE or agent session ids,
RUSHD_OUTPUT) are part of daemon identity, so each change restarts the daemon (confirmed). Excluding them from identity keeps one daemon, but then operations see no value, which breaks per-session attribution in build telemetry. Suggested fix: a per-request environment overlay for operations, so these variables are neither identity nor dropped.
Agent output and tests
- Agent-mode failure output drops the cause (confirmed). The failure summary drops the error line and the log path. A warnings-only build prints FAILURE with no project name, warning text or log path.
- The
rush-cli-clienttest suite isn't hermetic. WithCOPILOT_CLIset, which any Copilot agent shell does, 7 tests fail because output switches to agent mode.
- Warm-set maintenance is O(N²) and blocks the event loop (confirmed).
More daemon and client items from the same run
A few more items came up after my two previous comments. Line numbers are at 982b33b, as before. "Confirmed" means a second person reproduced the item independently.
Startup and liveness
- One failed or slow start blocks the daemon for the workspace until
daemon stop --force(confirmed).- When the launcher exits before readiness, or the startup helper gives up, the
.pid.json.startingreservation is kept on purpose (DaemonStartup.ts:105-118). At 982b33b the helper gets only the client's remaining 15 s deadline (connectOrStartDaemon.ts:452), so any start slower than that, for example under load, ends this way. Raising the helper's timeout only moves the threshold: with 120 s, a simulated 130 s start still ends this way. - While that file exists,
tryConnectAsynccloses every connection (connectOrStartDaemon.ts:259-267), so the client never uses a daemon that became ready later. - Every later client waits out the full 15 s deadline in the reservation loop (
:133-145). Then it falls back to in-process Rush, which also takes the repo lock. daemon statussaysready(or "could not connect") and doesn't mention the reservation.daemon restartstops the healthy daemon before it fails.- The file holds only a random token (
DaemonStartup.ts:32-33), so a client can't tell a live helper from a dead one. - Suggested fixes:
- If a client gets hello/ping from a live
pid.jsonowner, let it clear the reservation under the start mutex. - Record the helper's pid and start time in the file, so a dead helper fails at once with the reset hint instead of costing 15 s per request.
- Show the reservation in
daemon status. - Check for it before
restartstops anything. - The error message also has a doubled
...
- If a client gets hello/ping from a live
- When the launcher exits before readiness, or the startup helper gives up, the
-
daemon statusreports "no daemon" while the daemon is only busy (seen by two people).- A status call that lands during a warm-set maintenance pass exits 1 after 5 s with "did not complete hello/ping readiness within 5000ms". In one run, 7 of 206 calls failed this way, all right after builds.
- A caller that checks status before acting treats this as "not running".
- Suggested fix: report a live but busy
pid.jsonowner as its own state. The O(N²) pass from my previous comment is what triggers this.
- The client doesn't check that the daemon is alive while a request runs.
- The agent renderer's 10 s heartbeat keeps printing "running" after the daemon is stopped (SIGSTOP) or wedged, so the caller waits until its own timeout.
- A keepalive would let the client report "no response for N s". Either a ping on the request connection or progress ticks from the daemon would work.
- Separately, the
stop --forcehint doesn't work when the owner PID is alive but not ready.
Locking
-
On Linux,
LockFile.tryAcquirecan give the same lock to two or three processes at once (confirmed)._tryAcquireMacOrLinux(LockFile.ts:421-590) reads the folder once and breaks equal-birthtime ties by pid.- Birthtimes are coarse: 6 of 8 contenders had identical values on XFS.
- So a contender that listed the folder before a tied, smaller-pid file appeared can win, and the smaller-pid contender wins too. The comment at
:547-557assumes every contender sees every other file. - Repro without Rush: 8 node processes call
tryAcquireat a barrier and hold the lock for 1.5 s. Three people wrote separate harnesses: 3 of 60 trials had more than one holder on XFS each time, and 1-2 of 60 on btrfs. In every double grant, the later winner had the smaller pid. - Ties are the norm, not the exception:
birthtime.getTime()is whole milliseconds, and 8 files created back to back got only 56 distinct birthtimes out of 400, on XFS and btrfs alike. - With
rush build: 1 of 10 traced runs had two holders. In an untraced run, one of two concurrent builds failed with ENOENT on a file another build was restoring. - The daemon's execution lease and its reload use the same lock.
- Suggested fixes:
- Take the lock with an atomic exclusive create (
open(..., 'wx'), orlink()to a fixed name), and keep the pid files only for stale detection. - At minimum, treat an equal-birthtime live file you didn't create as a loss.
- Read
/proc/<pid>/statinstead of spawningps. That narrows the window (under load, decisions came 0.3-2.4 s after the barrier) but doesn't close it.
- Take the lock with an atomic exclusive create (
-
A contender with a different
TZor locale takes a lock that is already held (confirmed; 5 of 5 in each variant; deterministic from the code).getProcessStartTime(LockFile.ts:71-80) storesps -p PID -o lstart. procps formats lstart with the TZ and the locale of the process that runsps:LC_ALL=Cgives "Sun Sep 27 17:15:08 2026",LC_TIME=de_DE.UTF-8gives " So Sep 27 17:15:08 2026", and another TZ shifts the hour.- A contender compares the stored value with its own
psoutput. With a different TZ,LANG,LC_TIMEorLC_ALL, the live holder's file looks stale, so the contender deletes it and takes the lock. - A non-English
LANGis common. A developer shell withLANG=de_DE.UTF-8double-grants every time against a Rush process started withLC_ALL=Corenv -i, such as a script, an IDE task, a service or a daemon. - Suggested fixes:
- Run
pswithTZremoved andLC_ALL=C. In the common case (TZ unset, English or C locale) the stored bytes don't change, so mixed versions keep recognizing each other's locks. Switching the stored format (UTC or/proc/<pid>/stat) would break that. - Sturdier: when the strings differ, don't delete a lock whose pid is alive and whose process started no later than the lock file's birthtime (with slack for 1 s granularity). A reused pid always starts after the file was created.
- Run
-
The daemon's warm-set maintenance pass holds the repo lock (reproduced by four people).
#maintainOnceAsynctakes the execution lease (WorkspaceWarmSet.ts:289) and keeps it across#evictIdleAsyncand#reconcileWatcherPolicyAsync(:293-294). The lease is the samerushlock that in-process Rush takes (PhasedCommandEngineExecution.ts:31).- In a repo of about 1,800 projects the pass held it for 3-7 s after every request and about every 35 s while idle. A plain
rushcommand that reaches its lock check in that window fails with "Another Rush command is already running in this repository", though only the idle daemon is running: 6 of 6 and 4 of 4 times right after a daemon build, and about 1 in 7 while idle. In a repo of about 200 projects the same pass held it for 0.05-0.1 s, so the window grows with the pass's cost. - Making the pass fast (the O(N²) item in my previous comment) shrinks the window to milliseconds: with that fix, 4 of 4 plain
rush buildcommands right after a daemon build succeeded. Taking the lease only around an actual eviction would remove it. The message could also name the holder: a live daemon for this workspace, retry in a few seconds. That still matters for a native command started while a daemon build runs.
-
Every lock acquisition spawns
psfor the caller's own pid (one repro)._tryAcquireMacOrLinuxgets its own start time throughgetProcessStartTime(LockFile.ts:433), which runsps -p <pid> -o lstartsynchronously (:85). procps scans all of/proceven with-p, so on a host with about 12,500 processes each call took 0.26-0.30 s.- In-process Rush pays that once per command. The daemon takes the lease for every request, every maintenance pass and every graph request, so it pays it each time, on its event loop. In one measurement the lease took 0.29-0.42 s per request, most of a hot no-op request's time.
- Suggested fix: memoize the current process's start time; it never changes. The lock file content stays byte-identical, so mixed versions still recognize each other's locks, and other pids are still checked live.
Admission and graph loading
-
The false admission timeout from my previous comment is now confirmed. When it's fixed, stop counting graph-load time against configured timeouts too, not only the default one. Otherwise raising the timeout makes a request fail where the default would have succeeded, for example a build that arrives during a 90 s cold load with a 60 s timeout.
-
The admission hints recommend a variable that older Rush versions reject.
ClientAdmissionControls.ts:90,WorkspaceRequestAdmission.ts:233andWorkspaceRestartArbiter.ts:117all say to setRUSH_DAEMON_QUEUE_TIMEOUT_SECONDS.- In a repo whose
rushVersionpredates the daemon, exporting it makes every plainrushcommand fail with "not recognized by this version of Rush". The client's own in-process fallback strips it (launchClient.ts:275-277), so only plainrushbreaks, for example the nextrush install. - Suggested fix: while the selected version predates the daemon, suggest only
--wait-timeout <seconds>or the inline formRUSH_DAEMON_QUEUE_TIMEOUT_SECONDS=<s> rush-client ....
-
If any project's config can't be loaded, every daemon build in the workspace fails, while in-process Rush builds whatever doesn't need that project (confirmed).
- Triggers seen: a filtered install (
rush install --to <project>) leaves other projects' rig packages missing; a project'sconfig/rig.jsonnames a rig that isn't installed; a project'srush-project.jsonis invalid JSON. rush-client build -o <project>then fails in 1-2 s withroutingFailedand the load error, for exampleCannot find module '<rig package>/package.json' from '<a project that wasn't selected>'. That happens even when the broken project is outside the requested closure.--no-daemonbuilds the same target.- The engine loads configuration for every project (
PhasedScriptAction.ts:591,:693-703), and so does the workspace fingerprint (WorkspaceInputFingerprint.ts:281-282). In-process Rush loads it only for the selected closure. - There is no fallback: a
rejectedoutcome throws (launchClient.ts:260-262), and onlyfallbackreaches in-process Rush. The rejection text starts with unrelated terminal lines, and the cause is the last line. daemon statusstaysreadywith the graph not initialized, and the daemon log has no error. It isn't sticky: once the config is fixed, the same daemon builds again.- In a shared worktree, one person's half-edited
rush-project.jsonstops everyone else's daemon builds. - Suggested fixes: load project config lazily for the selected closure, or mark projects that can't load as unavailable instead of failing the graph. At minimum, fall back to in-process Rush with a one-line reason, and log the error.
- Triggers seen: a filtered install (
-
A rejected request on a fresh daemon prints every debug line from the graph load (confirmed).
EngineTerminalProvider.writebuffers messages of every severity while no request is executing (EngineTerminalProvider.ts:17-20), anddescribeErrorjoins all of them into the rejection text (:26-31, used atProductionDaemonRequestResolver.ts:116).- In a ~1.8k-project repo where most projects have no
rush-project.jsonof their own, a mistyped--toname on a fresh daemon printed 3,496 lines (583 KB): one "does not exist. Attempting to load via rig" debug line per project, with the one useful line at the very end. In-process Rush prints those lines only with--debug. - The same buffering puts unrelated lines ahead of the cause in the config-load failure above.
- The graph from a rejected request isn't kept, so every later rejection reloads it and repeats the output (6-14 s each) until a request succeeds.
- Suggested fixes: leave debug and verbose lines out of the buffer, or at least out of
describeError; cap a rejection message in the client and always print its last line; keep the loaded graph when only the selection fails.
Correctness
-
Update to the
:incrementalitem from my previous comment: it reproduces in a large repo with real Heft, and fixing the warm-set limits makes it reachable by default.- After a source file is deleted, the deleted file's outputs (
.js,.d.tsand maps) stay inlib/and go into the cache entry, under the same key in-process Rush computes. A later in-processrush buildof that state restores them. - With the default warm-set budget, a large repo evicts every retained result after each request, so
lastStateis gone and the initial--cleancommand runs. That hides the bug. Any change that keeps results (the limits item below) exposes it on every default daemon, so the two should land together. - Only one fix gives exact parity: a daemon
buildruns the initial command, as in-process Rush does, and watch mode keeps:incrementalwith cache writes off. Skipping the cache write after an incremental run stops the poisoning but leaves the stale files in the worktree.
- After a source file is deleted, the deleted file's outputs (
-
dependsOnEnvVarsis hashed with the daemon's own environment (confirmed).ProjectChangeAnalyzer.ts:455createsInputsSnapshotwithout anenvironment, so it defaults to the daemon'sprocess.env(InputsSnapshot.ts:93-96,:290). The declared variables are hashed from that (:401-406).- So if a request changes a variable that is excluded from daemon identity, the result is a wrong "no op" with stale output.
- After an in-process build has written the correct output, the daemon's next request restores its own stale cache entry over it.
- This interacts with the per-request environment overlay suggested in my previous comment. With the overlay alone, operations would see the request's value while the cache key uses the daemon's, and the cache would get entries labelled with the wrong key. Both have to change together: pass the request environment to
InputsSnapshot, or never exclude a variable that any project lists independsOnEnvVars.
-
A build that joins a batch while the batch reconciles can run on repo state from before its own edit (confirmed end to end by four people, 12 of 12 runs, on small repos and on a large one).
#executeBatchAsynctakes the execution lease and then awaitsreconcileInvalidationsAsync()(PhasedRequestRouter.ts:406-407). During that await the batch still accepts compatible joiners (#canJoinCurrentBatch,:355-367), which are taken at:410, and the iteration runs on the snapshot the reconcile read.- So a client that saves a file and then starts a build after the reconcile has read the repo gets a "no op" SUCCESS with stale output. The window is the rest of the reconcile, which in a large repo is mostly
git status. - The next build picks the edit up. But with several agents on one daemon, "save, then build while another agent's build is preparing" is the normal case: under load a reconcile took 20-50 s in the large repo.
- If the joined operation actually runs, its cache entry is written under the pre-edit key with outputs built from the edited file. A later build of the pre-edit state, for example after reverting the edit, restores the wrong outputs with rc 0, through the daemon and in-process alike. So a regression test for this should check the cache, not just the "no op".
- Suggested fixes: close the batch before reconciling (take the compatible pending requests, stop accepting, then reconcile), or admit only requests enqueued before the reconcile started.
-
The input-change guard from [rush-lib] Skip build cache writes when an operation's inputs changed during execution #6086 misses a file saved while the repo state is being read, in-process too (one repro, daemon and in-process).
captureInputFilesStateruns inbeforeExecuteIterationAsync(CacheableOperationPlugin.ts:171,:225), aftergetRepoState'sgit status. A tracked file saved between the two (for example duringgit hash-object, which in a large repo takes seconds) is already in the captured stats, sohaveInputFilesChanged(:595) sees no change.- The operation builds the new content, and the entry is written under the key of the old content. Reverting the file and building again restores the wrong outputs from the cache with rc 0.
- Suggested fix: record a timestamp before
git status, and inhaveInputFilesChangedtreat any input whose mtime is at or after it as changed. The capture already stats every input, so this is one comparison per file.
-
After any daemon request that executes an operation, the operations outside its selection lose their trust, and their consumers stop writing build-cache entries (found in a repo of about 200 projects and confirmed by three more people, each on their own small repo; in-process Rush on the same engine writes every entry). Outputs stay correct; the cache entries are what's lost.
- For each request the router disables every operation in the graph, then enables the request's selection (
PhasedRequestRouter.ts:1015-1026). The iteration still makes a record for every operation, so the unselected ones finish Skipped withenabled === false. - The Skipped branch treats that like any other skip: it blocks cache writes for consumers and deletes the operation's trusted state hash (
CacheableOperationPlugin.ts:640-647). A skip leaves the retained result in place (OperationGraph.ts:1391-1399), so in later iterations the operation is unchanged and skipped again (PhasedOperationPlugin.ts:185-193), and every consumer that rebuilds runs with writes blocked. It recovers only when that operation runs or restores again: an input change, a graph reload, or an in-process build in between. --only Xof a project with dependencies is the simplest trigger, because X itself then runs with writes blocked, as in-process Rush does. But--onlyisn't needed: after--to Aexecutes anything, a later--to Bthat rebuilds a consumer of a project outside A's closure writes no entry for it. A request that executes nothing doesn't trigger it.- With several people or agents on one daemon, almost every executed request selects a small closure, so the next request's rebuilt projects write no entries. Other worktrees, CI and in-process builds lose those entries.
- Suggested fix: in the Skipped branch, treat an operation that is disabled only because the request didn't select it like a retained skip. Block its consumers only when its trusted hash differs from the current state hash, and don't delete its trust. That also lets
--only Xwrite entries when X's dependencies are trusted.
- For each request the router disables every operation in the graph, then enables the request's selection (
Scale
- With the default warm-set limits, a hot request in a large closure is slower than in-process Rush restoring everything from the cache (confirmed).
- For a closure of about 790 projects and 1,576 operations, the default budget (512 MB, 20 projects) is below the daemon's own RSS (0.65-1.7 GB), so each maintenance pass evicts every retained result.
- Every hot request then restores 772 operations from the cache: 58-68 s, against 42-56 s for in-process Rush doing the same restores.
- Releasing only retained results that hold resources, and keeping the rest, made the same hot request 6.3-7.1 s, all 1,576 operations no-op.
- Concurrent requests appear to be prepared one after another.
- Hot, disjoint clients were started together in bursts. The time until the burst's single
git statusgrew by about 1.9 s per extra client: 5.6 s at N=1 and 18.7 s at N=8. - That matches the per-request
rush-project.jsoncost from my previous comment, so fixing that item should help bursts the most. - The burst itself is handled well:
- one snapshot and one execution, with correct results;
- 8 clients finish about 3× sooner than running them one after another;
- in a ~200-project repo it is 8× hot and 6× with edits.
- 4 clients whose first builds overlapped ran each operation once, each client finished when its own selection did, and the outputs matched a clean in-process rebuild.
- Hot, disjoint clients were started together in bursts. The time until the burst's single
Plugins (addendum to the plugin list)
- Plugins have no contract for detecting that they run inside the daemon engine.
onExecuteAsyncreturns early in engine mode (RushCommandLineParser.ts:418-421). Reading the code,parser.telemetryis then never created, sobeforeLognever fires for daemon iterations.- A plugin that starts per-iteration work in
beforeExecuteIterationAsyncand stops it inbeforeLogtherefore leaks one timer per iteration, and each timer keeps the old graph alive. - Today the only signal is
generateFullGraph === true && !isWatchon the graph context (PhasedScriptAction.ts:591), and that isn't a contract. - An explicit engine flag would fix this, together with the per-request telemetry and request-scoped hook items from my plugin comment.
- One failed or slow start blocks the daemon for the workspace until
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNeeds triage
Re-architect Rush from a CLI-first tool into a server-first system: a long-lived per-workspace
daemon (
rushd) owns the engine and warm build state; thin, presentation-owning clients (CLI first;MCP/web later) talk a presentation-free protocol. Delivered as new split packages merged
incrementally into
mainat0.x, culminating in a Rush 6 cutover.Design rationale & discussion: #5863 (RFC).
Hard dependencies
IOperationGraph(
setEnabledStates,scheduleIterationAsync,invalidateOperations,abortCurrentIterationAsync,closeRunnersAsync) andOperationGraphHooks(onIdle,onExecutionStatesUpdated).its "daemon-aligned major default flip" is the Rush 6 cutover.
Workstreams (sub-issues)
Definition of done for the Epic
rush/rushxdefault to the daemon with a permanent in-process fallback (--no-daemon/RUSH_DAEMON=0).6.0.0in the same major as [rush] Rush Reporter Overhaul #5858 phase 6.