Restart a workspace whose loop counter is exhausted - #640
Conversation
|
LGTM! |
Preview:
|
|
LGTM! |
The Workers runtime refuses a Durable Object call once the loop
counter behind it is spent ("Subrequest depth limit exceeded. This
request looped back into the Workers runtime too many times."). A
workspace object's outgoing channels can end up holding a spent
counter with nothing recursing, and from then on every call it makes
to a user object is refused until the instance is replaced.
When one of the workspace's own user-object calls is rejected that way
(a call through a wrapped stub, the last-active bump, or the outputs
sync), the workspace now schedules the existing access restart. At
most one restart per instance, and none in an instance's first 60
seconds. Errors thrown by gadget, agent or gatekeeper-facet code are
never consulted.
The shared test fixture gains the new helper so existing suites keep
passing.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Covers the message predicate (the sibling too-many-stages error, text quoted inside another error and non-Error values do not match), the wrapper's onRejection hook, and the restart itself through the last-active bump, the outputs sync and wrapped owner and session stubs, including the 60 second floor and the one-per-instance limit. Co-Authored-By: Claude Code <noreply@anthropic.com>
The drain asks the initiator's user object for its chat context on a bare stub, so a loop-limit rejection there was logged (agent.callback.start.failed) but did not restart the workspace. The stub now goes through wrapUserDo like the workspace's other watched user-object calls. The drain itself is unchanged: the rejection is still logged, the calls stay recorded and the alarm retries. Co-Authored-By: Claude Code <noreply@anthropic.com>
A recorded call whose drain is refused restarts the workspace, still logs the drain's own failure and stays recorded for the retry. Co-Authored-By: Claude Code <noreply@anthropic.com>
e7cc8b5 to
f0c6097
Compare
|
LGTM! |
Co-Authored-By: Claude Code <noreply@anthropic.com>
|
LGTM! |
Eval resultsVerdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.
Failed checks
|
🔬 Eval runs reviewPerformanceChange-calendar, chess and worker-logs stayed at 10/10 passes; incident-desk fell from 10/10 to 9/10, within noise according to ⚪ VERDICT: NO REGRESSION FROM THIS PRThe lone failure is a generated-code mistake unrelated to the diff’s loop-limit recovery in TriageFailure modes
Tool errors
What to do
|
Currently a workspace can get stuck/borked with every action failing on
The workspace stays stuck until the overseer DO is restarted, which until now only happened on a redeploy or when the object was evicted. This is because a dynamic worker is pinned to the loop counter of whichever request was current when it was created. Each call it makes back into the workspace arrives one lower and becomes the current request, so the next dynamic worker is pinned lower again, until the counter reaches zero.
This is a bandaid fix: the workspace now restarts itself when one of its own calls to a user object is rejected with that error (
isLoopLimitError). When this happens:OverseerImpl.restartIfLoopLimitedcalls the existingscheduleAccessRestart, at most once per instance and never in an instance's first 60 seconds.wrapDoStubForTelemetryjust gains an optionalonRejection, and the nine call sites now go throughOverseerImpl.wrapUserDo)workspace.loop.limit.restartat error level.Known limits:
Related: #282 (possibly the same runtime behaviour).