Skip to content

Restart a workspace whose loop counter is exhausted - #640

Merged
Maximo-Guk merged 5 commits into
mainfrom
maximo/loop-limit-self-heal
Oct 2, 2026
Merged

Maximo-Guk merged 5 commits into
mainfrom
maximo/loop-limit-self-heal

Conversation

@Maximo-Guk

@Maximo-Guk Maximo-Guk commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Currently a workspace can get stuck/borked with every action failing on

Subrequest depth limit exceeded. This request looped back into the Workers runtime too many times.

The workspace stays stuck until the overseer DO is restarted, which until now only happened on a redeploy or when the object was evicted. This is because a dynamic worker is pinned to the loop counter of whichever request was current when it was created. Each call it makes back into the workspace arrives one lower and becomes the current request, so the next dynamic worker is pinned lower again, until the counter reaches zero.

This is a bandaid fix: the workspace now restarts itself when one of its own calls to a user object is rejected with that error (isLoopLimitError). When this happens:

  • OverseerImpl.restartIfLoopLimited calls the existing scheduleAccessRestart, at most once per instance and never in an instance's first 60 seconds.
  • It is uses the existing wrapped user-object stubs (wrapDoStubForTelemetry just gains an optional onRejection, and the nine call sites now go through OverseerImpl.wrapUserDo)
  • Each restart logs workspace.loop.limit.restart at error level.

Known limits:

  • The call that hits the error still fails, so the user sees one failure before the restart. An in-flight agent turn is resumed as after any restart.
  • A gadget whose own calls fail while the workspace's calls still succeed is not detected.
  • Proper fix is obviously in the runtime

Related: #282 (possibly the same runtime behaviour).


Devin Review

@github-actions github-actions Bot added the kernel Changes to the Workshop kernel label Oct 2, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Preview: pr640-maximo-loop-l-d0832e49

https://pr640-maximo-loop-l-d0832e49-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

@Maximo-Guk
Maximo-Guk marked this pull request as ready for review October 2, 2026 16:26
devin-ai-integration[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
Maximo-Guk and others added 4 commits October 2, 2026 16:26
The Workers runtime refuses a Durable Object call once the loop
counter behind it is spent ("Subrequest depth limit exceeded. This
request looped back into the Workers runtime too many times."). A
workspace object's outgoing channels can end up holding a spent
counter with nothing recursing, and from then on every call it makes
to a user object is refused until the instance is replaced.

When one of the workspace's own user-object calls is rejected that way
(a call through a wrapped stub, the last-active bump, or the outputs
sync), the workspace now schedules the existing access restart. At
most one restart per instance, and none in an instance's first 60
seconds. Errors thrown by gadget, agent or gatekeeper-facet code are
never consulted.

The shared test fixture gains the new helper so existing suites keep
passing.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Covers the message predicate (the sibling too-many-stages error, text
quoted inside another error and non-Error values do not match), the
wrapper's onRejection hook, and the restart itself through the
last-active bump, the outputs sync and wrapped owner and session
stubs, including the 60 second floor and the one-per-instance limit.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The drain asks the initiator's user object for its chat context on a
bare stub, so a loop-limit rejection there was logged
(agent.callback.start.failed) but did not restart the workspace. The
stub now goes through wrapUserDo like the workspace's other watched
user-object calls. The drain itself is unchanged: the rejection is
still logged, the calls stay recorded and the alarm retries.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A recorded call whose drain is refused restarts the workspace, still
logs the drain's own failure and stays recorded for the retry.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@Maximo-Guk
Maximo-Guk force-pushed the maximo/loop-limit-self-heal branch from e7cc8b5 to f0c6097 Compare October 2, 2026 21:32
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
Co-Authored-By: Claude Code <noreply@anthropic.com>
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Eval results

Verdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.

Task Score Δ score Fisher test Cache hits Avg min Avg steps
change-calendar 100% 0 pp p = 1.00 87% → 88%
+1 pp
4.0 → 3.2 27.0 → 26.0
chess 100% 0 pp p = 1.00 97%
0 pp
9.7 → 7.9 58.2 → 65.2
incident-desk 100% → 90% −10 pp p = 1.00 95% → 94%
−1 pp
4.6 → 3.6 34.8 → 28.9
worker-logs 100% 0 pp p = 1.00 92% → 93%
+1 pp
4.0 → 3.6 25.0 → 26.0
Failed checks
Task Check Failed
incident-desk t1 opens-acknowledges-and-resolves-in-order 0/10 → 1/10
incident-desk t1 simultaneous-acknowledges-yield-exactly-one-owner 0/10 → 1/10

Run · trajectories and raw results

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

🔬 Eval runs review

Performance

Change-calendar, chess and worker-logs stayed at 10/10 passes; incident-desk fell from 10/10 to 9/10, within noise according to comparison.json, and no prompt-cache change was significant. Candidate runs were faster on every task, but costs overlapped main: mean chess cost rose from approximately $0.058 to $0.062 while calendar and incident-desk became slightly cheaper, with no demonstrated connection to this diff. Chess consumed the most steps, rising from 58.2 to 65.2 per run; candidate trial 7 wasted 14 editFile calls on mismatched text, while incident-desk trial 4 accounted for the only failed run.

⚪ VERDICT: NO REGRESSION FROM THIS PR

The lone failure is a generated-code mistake unrelated to the diff’s loop-limit recovery in do-retry.ts and overseer.ts; neither side’s trajectories shows the loop-limit rejection needed to demonstrate that recovery’s benefit.

Triage

Failure modes

  • Losing acknowledgements throw instead of reporting the owner · incident-desk 1/10 · model error · this PR: no — In trial 4, turn 1’s writeFile(server.js) used .one() on a conditional UPDATE … RETURNING that returns no rows after someone else claims the incident, so repeated and simultaneous acknowledgements threw before reaching the ALREADY_ACKNOWLEDGED branch. The final reply nevertheless claimed later callers would learn who won; main’s passing implementations handled unsuccessful claims explicitly, and this PR changes neither gadget SQL semantics nor the task.

Tool errors

  • editFile: No matching text was found in the file. · chess 6 → 17, incident-desk 0 → 2, worker-logs 2 → 4 · model error — Chess trial 7, turn 2 repeatedly supplied incorrectly escaped strings and an extra space absent from the source, even after grep showed the exact text; 14 failures were concentrated in that run.
  • editFile: Validation failed for tool "editFile": · change-calendar 2 → 4, chess 5 → 3, incident-desk 9 → 2, worker-logs 4 → 1 · model error — Agents omitted required arguments despite the schema; for example, calendar trial 3, turn 3 omitted filename, then successfully retried with client.js after the error identified the missing field.
  • editFile: Multiple matches were found. The text to match must be unique. · change-calendar 2 → 0, chess 1 → 3, incident-desk 7 → 4, worker-logs 1 → 1 · model error — Agents supplied repeated fragments rather than unique context. Main’s incident-desk trial 1 recovered by including method-specific context; candidate chess trial 5, turn 3 retried the same ambiguous fragment once before expanding it.
  • readFile: File does not exist. · change-calendar 2 → 6, worker-logs 2 → 4 · model error — Agents inspected server.js and client.js immediately after creating a blank gadget, then wrote them; calendar candidate trial 2 shows both reads failing despite the unchanged system prompt explicitly saying new gadgets have no files.

What to do

  • No eval-driven change is required for this PR. Optional, separate hardening in packages/workshop-backend/src/agent.ts: add a storage reminder that SQL .one() throws on zero rows and optional matches, including unsuccessful conditional updates, should use .toArray()[0] instead.

github run

@Maximo-Guk
Maximo-Guk merged commit 10df02c into main Oct 2, 2026
22 checks passed
@Maximo-Guk
Maximo-Guk deleted the maximo/loop-limit-self-heal branch October 2, 2026 22:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kernel Changes to the Workshop kernel

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants