Skip to content

Let gadget UIs import other modules - #642

Draft
ndisidore wants to merge 8 commits into
mainfrom
feat/multi-file-gadget-uis
Draft

ndisidore wants to merge 8 commits into
mainfrom
feat/multi-file-gadget-uis

Conversation

@ndisidore

@ndisidore ndisidore commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

client.js can now import the gadget's other .js files, as server.js already can, so a UI no longer has to fit in one 512K-character file. This work in done in such a way that, if, in the future, we pursue TypeScript and bundling it can arrive without gadget code changing.

  • Walk: find every module client.js reaches, so a use-role viewer gets nothing else (never server.js)
  • Rewrite: point each relative import at an internal gadget: key, since relative imports fail inside data: modules
  • Bundle: carry the extra modules and diagnostics beside the existing jsCode, so tabs opened before this change still load single-file gadgets
  • Page: load each module from its own data: URL through an import map, with sourceURL last so errors name the right file and line
  • Runtime: set up RPC and the gadget, RpcTarget and RpcStub globals before any gadget module runs, with nothing prepended to client.js
  • Diagnostics: report bad imports in the gadget console, where the agent sees them, and show a fatal one in place of the frame
image image

The live iframe's CSP is unchanged. The export page adds a per-render nonce so its import map runs under script-src data:. The agent prompt covers splitting files, the import rules, and that anyone who can use a gadget receives everything client.js reaches.

UiBundle gains optional modules and diagnostics. ui-page turns a bundle into
base64 data: URLs with a trailing sourceURL and an import map keyed in the
gadget: namespace, shared by the live iframe and browser export.
@github-actions github-actions Bot added workshop/frontend Changes to the Workshop frontend kernel Changes to the Workshop kernel workshop/shared Changes to shared Workshop APIs labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Preview: pr642-feat-multi-fi-5cb2fc55

https://pr642-feat-multi-fi-5cb2fc55-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

buildUiBundle walks from client.js with es-module-lexer's /js build,
rewrites each relative import to its gadget: key, and reports the import
rules' violations as diagnostics. A use-role viewer receives only what
client.js reaches, never server.js or modules only it imports.
The export page adds the bundle's import map, admitted by a per-render CSP
nonce rather than 'unsafe-inline', and its runtime sets gadget, RpcTarget
and RpcStub as globals instead of prepending them to client.js. Export
fails with the first fatal diagnostic before launching a browser.
The prelude becomes a runtime module script ahead of the entry, setting
gadget, RpcTarget and RpcStub as globals, so nothing is prepended to
client.js and its line numbers are exact. Modules load through the import
map; errors are forwarded with their file, line and column, including
SyntaxErrors and failed module loads; diagnostics reach the gadget
console, and a fatal one replaces the frame.
Covers the import rules, which ones apply only to client code, the
client globals, and that everything client.js reaches is sent to anyone
who can use the gadget.
* above, for caching. Or... maybe we should actually serve over RPC, but also employ the
* Cache API in the browser? Or some other local storage?
*/
jsCode: string;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What if there are cyclic imports such that another module imports client.js?

Maybe we should say: If there are multiple files, we don't send jsCode at all. We only send modules, plus maybe mainModule to name the entrypoint?

(Eventually we can phase out jsCode and always send modules, though probably only after clients have updated -- an update which totally breaks gadgets until everyone refreshes would be too disruptive.)

Comment thread packages/workshop-shared/src/api.ts Outdated
modules?: UiModule[];

/** Problems found while collecting the UI's modules. Absent when there are none. */
diagnostics?: UiDiagnostic[];

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This diagnostics representation strikes me as overcomplicated and YAGNI.

What if:

  • Dynamic imports aren't supported for now. What's even the point of a dynamic import when the code is all downloaded upfront anyway?
  • Any error just causes getUiBundle() itself to throw?

Then on the client side we catch the getUiBundle() error and provide a one-click button to report it to the agent.

window.parent.postMessage("handshake", "*", [port2]);
gadget = newMessagePortRpcSession(port1);
// RPC stub to the gadget's server-side Durable Object.
globalThis.gadget = newMessagePortRpcSession(port1);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've never really liked the hack of gadget simply being a global, and I like it less in a world of multiple modules and large applications that import npm libraries and so on. gadget is a capability. Some random npm library shouldn't just be able to grab it out of the global scope...

Relatedly, I want to introduce a second RPC stub that talks to the workshop UI, to do things like display a specific agent in the chat sidebar.

What if, when using modules, you have to export a main function:

export default {
  async function main(gadget, workshopUi) {
    ...
  }
}

And RpcStub and RpcTarget should of course be imported from "capnweb" instead of assumed to exist as globals? (We can special-case "capnweb" for now until we have proper npm imports.)

Comment thread packages/workshop-backend/src/browser-export.ts Outdated
Comment thread packages/workshop-backend/src/ui-bundle.ts Outdated
…dules

A multi-file UI's bundle is now {modules} with client.js first, and a
single-file one stays {jsCode}. buildUiBundle throws on the first import
that breaks the rules for UI code instead of returning diagnostics, so
UiDiagnostic, the warnings and export's separate fatal check are gone. A
bad string-literal import() still rejects only when it runs. The total
size cap is dropped: it fired well below any limit the bundle hits.
Matches the live iframe's policy instead of minting a per-render nonce.
data: scripts already let gadget code run anything in the page.
A failed load shows its message in place of the frame and sends it to the
gadget console, which the composer can attach for the agent. It counts as
loaded, so the agent's next code change reloads the view.
@ndisidore
ndisidore force-pushed the feat/multi-file-gadget-uis branch from cba84fa to 3d99557 Compare October 2, 2026 19:33
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Eval results

Verdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.

Task Score Δ score Fisher test Cache hits Avg min Avg steps
change-calendar 100% → 90% −10 pp p = 1.00 87% → 88%
+1 pp
4.0 → 3.8 27.0 → 24.2
chess 100% → 80% −20 pp p = 0.47 97% → 96%
−1 pp
9.7 → 7.3 58.2 → 47.4
incident-desk 100% 0 pp p = 1.00 95% → 94%
−1 pp
4.6 → 3.9 34.8 → 30.9
worker-logs 100% → 60% −40 pp p = 0.09 92% → 93%
+1 pp
4.0 → 2.9 25.0 → 21.8
Failed checks
Task Check Failed
change-calendar t1 agent.timedOut 0/10 → 1/10
chess t1 agrees-with-the-oracle-on-the-hard-positions 0/10 → 2/10
chess t1 agrees-with-the-oracle-on-perft-positions 0/10 → 2/10
chess t1 agrees-with-the-oracle-through-random-games 0/10 → 2/10
chess t1 starts-from-the-standard-position 0/10 → 1/10
chess t1 the-game-is-shared-across-connections 0/10 → 1/10
worker-logs t1 hourly-buckets-cover-every-hour-including-empty-ones 0/10 → 4/10
worker-logs t1 ranges-are-half-open-and-filter-by-worker 0/10 → 4/10
worker-logs t1 ingests-resets-and-summarises-per-worker 0/10 → 3/10

Run · trajectories and raw results

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

🔬 Eval runs review

Performance

Across 10 runs per side, pass rates went from 100% to 90% for change-calendar, 80% for chess, 100% for incident-desk, and 60% for worker-logs; comparison.json classifies every change as within noise. Mean run costs fell respectively from $0.0292 → $0.0252, $0.0582 → $0.0456, $0.0282 → $0.0245, and $0.0229 → $0.0166, with fewer steps on every task, but early failures truncate the candidate’s work and the differences overlap individual-run variation. Worker Logs lost the most runs, four at turn 1, while chess incurred the largest retry overhead, with unmatched edits rising from 6 to 26; cache hits barely moved, and significant cache-break changes were mixed rather than a consistent improvement.

⚪ VERDICT: NO REGRESSION FROM THIS PR

The observed failures do not follow from the diff’s multi-file UI support or import guidance, pass-rate changes remain within noise, and neither the lower costs nor the mixed caching changes establish an attributable improvement.

Triage

Failure modes

  • Oversized SQL inserts · worker-logs 2/10 · model error · this PR: no — In turn 1, trials 4 and 6 wrote server.js with 100-event inserts requiring 600–700 bound parameters; ingestion subsequently failed with “too many SQL variables,” leaving the later reporting checks without their data. Trial 4’s executeCode checked only three events and missed the scale problem; main’s passing implementations included per-event inserts. The diff does not change SQL execution.
  • Hourly aggregation misses populated buckets · worker-logs 1/10 · model error · this PR: no — Trial 2’s turn-1 writeFile(server.js) grouped timestamps using division without explicitly flooring them, then looked up UTC hour-start keys; hourly() returned zero counts despite successful ingestion and correct summaries. Main trial 1 explicitly floored timestamps before grouping; neither implementation was directed by the new import guidance.
  • Unsupported SQL transaction statements · worker-logs 1/10 · model error · this PR: no — Trial 7’s turn-1 writeFile(server.js) used BEGIN TRANSACTION; the runtime rejected ingestion and explicitly requested the storage transaction APIs instead. The agent declared completion without exercising ingestion, and this PR leaves that runtime restriction unchanged.
  • Malformed client JavaScript · chess 1/10 · model error · this PR: no — Trial 2’s turn-1 writeFile(client.js) omitted a closing template-literal delimiter in the piece-rendering code. A later editFile fixed an unrelated DOM append, but the completion reply claimed success without execution; verification could not start the Worker. This was single-file generated code, not a rewritten import.
  • Sliding-piece attacks checked only at distance one · chess 1/10 · model error · this PR: no — Trial 8’s turn-1 writeFile(server.js) restricted rook and bishop attack detection to adjacent squares, allowing moves and castling through distant attacks. Subsequent edits addressed king captures and mutation serialization, not that rule; the final reply nevertheless claimed full rules support.
  • Agent activation stalls before completion · change-calendar 1/10 · harness bug · this PR: no — Trial 7 stopped after turn-1 createGadget and two successful writeFile calls, then hit the 420-second activation timeout without a reply or tool error. The transcript does not localize the stall, but shows no import-related operation, and the diff does not change activation or timeout handling.

Tool errors

  • editFile: No matching text was found · change-calendar 0 → 2, chess 6 → 26, incident-desk 0 → 1, worker-logs 2 → 2 · model error — Agents supplied text differing from the file’s escaping or whitespace. Chess trial 3 repeatedly retried the same mismatching PGN replacement instead of using the exact current text; the unchanged tool contract requires an exact match.
  • editFile: Validation failed · change-calendar 2 → 3, chess 5 → 1, incident-desk 9 → 4, worker-logs 4 → 4 · model error — Calls omitted required arguments, notably filename; incident-desk trial 1 and worker-logs trial 6 recovered by resubmitting with the filename.
  • editFile: Multiple matches were found · change-calendar 2 → 1, chess 1 → 5, incident-desk 7 → 2, worker-logs 1 → 0 · model error — Agents attempted global-style replacements with short repeated snippets. Candidate chess trial 1 tried replacing several repeated method calls despite the tool’s explicit exactly-one-location requirement; main showed the same mistake.
  • readFile: File does not exist · change-calendar 2 → 0, incident-desk 0 → 4, worker-logs 2 → 0 · model error — Agents read server.js and client.js immediately after creating an empty gadget, before writing them. The system prompt already explains that non-blueprint gadgets start without files.

Prompt cache

  • change-calendar · cache hits 86.9% → 87.9% · breaks 7.69% → 6.29% · this PR: no — Trial 1 on both sides loses the history cache at the starts of turns 2 and 3 as the gadget/file inventory changes, including the newly created maintenance document. Cache reads fall to the static prefix, 6,948 → 7,169 tokens; the extra 221 tokens are the new prompt guidance, not fewer misses. Mean steps differ, 27.0 → 24.2, and the timeout skips later turns.
  • incident-desk · cache hits 94.5% → 93.6% · breaks 0.98% → 1.28% · this PR: no — Both trial-1 transcripts rebuild the history suffix at turn 2 after gadget creation, reading only the same static prefix, again 6,948 → 7,169 tokens. Turn 3 retains history caching on both sides; no additional candidate miss stands out. Fewer steps, 34.8 → 30.9, change the rate’s weighting.
  • worker-logs · cache hits 92.4% → 92.8% · breaks 1.63% → 1.20% · this PR: no — Passing trial 1 has the same turn-2 inventory-induced break and retains history reads at turn 3 on both sides. Four candidate failures never reach turn 2, avoiding that later miss rather than improving caching; mean steps also fall from 25.0 → 21.8.

What to do

  • Optional: add a completion-check sentence to packages/workshop-backend/src/agent.ts asking agents to exercise representative RPC inputs, including realistic ingestion batches and populated hourly buckets, before claiming success; these failures do not require changing this PR’s module support.
  • Optional: add recovery guidance to EDIT_FILE_TOOL_DESCRIPTION in packages/workshop-backend/src/agent.ts: after a failed match, reread the relevant text and use unique context rather than retrying an unchanged replacement.

github run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kernel Changes to the Workshop kernel workshop/frontend Changes to the Workshop frontend workshop/shared Changes to shared Workshop APIs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants