English | 简体中文
The Web UI (packages/webui) is the browser frontend for MiniMax Code. It uses the same engine as the TUI — the CLI's ACP server (mcode acp, JSON-RPC 2.0 over stdio) — so terminal, browser, and desktop clients run against one runtime. Each turn runs over ACP by default; two conditions switch it to the one-shot mcode exec CLI (see Transport selection). It is not a plugin: it ships inside the repository and is launched by the CLI.
This document describes what the shipped webui does today, against the source tree. Every claim links to the file or test that backs it. When something is partial or a placeholder, the row below says so plainly. The shape and limits recorded here come from packages/webui/{server,webapp} as of the file-level ages noted in each subsection.
mcode-web # http://127.0.0.1:18090
mcode web # equivalent — `web` and `webui` both resolve
mcode webui --port 8123 --host 127.0.0.1
mcode webui --token "$(openssl rand -hex 16)" --host 0.0.0.0 # LAN, token-gated
pnpm mcode-web # from a source checkout (built)
node packages/webui/server.js # direct, from a checkoutThe command resolves the webui package (installed dist/webui/ or source packages/webui/), spawns the server as a child process, and points it back at the running CLI through MCODE_WEBUI_SELF_ENTRY. The webui then spawns node <cli> acp per active browser tab.
Without --port the server starts on 18090 and moves to the next free port when 18090 is taken, logging the URL it bound — the launcher opens that one. An explicit --port (or PORT) is pinned: it never moves, so a taken port exits with EADDRINUSE instead. (See packages/webui/server/lib/config.js#PORT and the launcher in packages/webui/server.js.)
The development Web UI runs alongside an installed official mcode without
conflicts: the dev webui always spawns the checkout's own dist/cli.js as its
engine (detection order: MCODE_CMD > MCODE_WEBUI_SELF_ENTRY > repo
dist/cli.js > ~/.minimax-code > PATH), and it shares the host's
~/.minimax sessions and ~/.mcode-webui state.
From a checkout of this repository:
corepack pnpm install && corepack pnpm build # once, and after engine changes
node dist/cli.js webui # dev Web UI on 127.0.0.1:18090
node dist/cli.js webui --port 8123 # keep the installed one freeWhen iterating on the Next.js frontend in packages/webui/webapp/ you want both
the Node backend (port 18090, serves /api/*) and the Next dev server (port
18091, with HMR, proxies /api/* → 18090) running at once. pnpm run webui:dev
boots both in a single shell, prefixes their output so you can tell which side
is talking, and tears them down together on Ctrl+C:
pnpm run webui:dev # http://127.0.0.1:18091/ ← open this in the browserIt is a thin wrapper over node scripts/dev-webui.mjs with no extra
dependencies. Stop the official mcode runtime first if port 18090 is busy
(pkill -f "dist/cli.js webui"), or pass --port 28090 to mcode webui and
export MCODE_WEBUI_ORIGIN=http://127.0.0.1:28090 so the dev proxy targets
the right backend.
| Command | What it does |
|---|---|
pnpm run webui:dev |
Start both the backend (:18090) and Next dev (:18091, HMR) together; Ctrl+C cleans up both. |
pnpm run webui:build |
next build the webapp (packages/webui/webapp/out/ is the static export). |
pnpm run webui:start |
Serve the already-built webui via node packages/webui/server.js on :18090 (no HMR). |
pnpm run webui:typecheck |
tsc --noEmit over the webapp's TS sources. |
pnpm run webui:test |
Run all webui unit tests — backend (test:webui) + frontend (test:webapp). |
next start is intentionally omitted: the webui ships as a next export static
build and the backend serves those files directly, so there is no Next server
runtime to start. ESLint is also not wired into the webapp yet — add it via
npx next lint once a .eslintrc is in place.
The repository's Docker setup runs the branch in a clean environment: no host home directories are mounted, so host models/tools/sessions never leak in (and container state never leaks out). Model credentials come from the environment — each collaborator tests with their own key:
MINIMAX_CN_API_KEY=... docker compose up webui # MiniMax cn region
MINIMAX_API_KEY=... docker compose up webui # MiniMax global region
# open http://localhost:18080/?token=dev-token
docker compose down # reset to factory stateThe image runs as a non-root user (uid 1000) with a realistic home at /home/user (Desktop/Documents/Downloads/Pictures/Music/Videos/projects + XDG config), so the directory picker's well-known-folder keywords behave as on a desktop. docker/entrypoint.sh seeds a fresh in-container ~/.minimax/config.yaml
(minimaxModelSource: minimax_api_key + the key + minimax_api/MiniMax-M3
as the default model); MAVIS_REGION is derived from which variable you set
and can be overridden explicitly. Because the host browser is a non-local
client from the container's perspective, every URL carries ?token=…
(WEBUI_TOKEN, default dev-token); the port is WEBUI_PORT (default
18080). A container without either key variable starts fine but chat has no
model credentials until one is provided.
For interactive development over the mounted source (same env-key flow):
docker compose run --rm -p 18080:18080 dev
# inside the container:
pnpm install --no-frozen-lockfile && pnpm build
node dist/cli.js webui --host 0.0.0.0 --no-open # PORT defaults to 18080- Loopback bind by default; LAN exposure requires
--host/HOSTenv or the persistedlanBindsetting. - Trusted-origin CORS + a browser Origin/CSRF gate that applies even to loopback requests.
- Token auth (
?token=/Authorization: Bearer) for non-local requests; local requests bypass. - Read-only mode for non-local sessions; per-request
authorize()gate with fail-closed audit; rate limiting; workspace containment; bounded uploads; no telemetry. - Credential-shaped file previews are refused by default (slice 16). A basename match against
.env/.env.*,*.pem/*.key,id_rsa/id_ed25519/id_ecdsa/id_dsa,known_hosts,authorized_keys,.npmrc,.pypirc,.netrc,.pgpass,credentials*, plus the backup-suffix set (.bak/.old/.orig/.backup/.save/.swp) returns HTTP403 {code: "credential"}fromGET /api/fs/read-file. The webapp renders a "仍要打开?" second-confirmation; reopening the same URL with?confirm=1gets the plaintext. The shared predicate lives inpackages/webui/server/lib/credential-file.jsand is mirrored verbatim inpackages/webui/webapp/lib/credential-file.ts; the test suite (packages/webui/webapp/test/credential-file.test.ts) walks both implementations on the same fixtures so they cannot drift. Tree listing,/api/fs/search, and OS-default open/reveal do not bypass the gate, but they are not plaintext previews and remain available — search flags matches (credential: true) but never returns contents. The predicate is name-based and therefore does not defend against hardlink aliasing (two names that share an inode — e.g.config.txt → .env— are indistinguishable by basename; the kernel does not expose the "primary" name from the inode). The defence covers symlinks (resolved byrealpathSync) but not hardlinks — operators concerned about hardlink aliasing must keep the workspace tree uncluttered.
The canonical disclosure is packages/webui/references/SECURITY-NOTES.md.
Slice 27 extends the same posture to the write side: POST /api/fs/write
runs the identical containment gate, applies the identical credential
predicate (default-refuse; confirm:true releases the write and emits a
credential.override audit line with endpoint:"write"), and
conflict-checks the caller's (mtime, size) baseline so an external edit
surfaces as 409 {code:"conflict"} instead of being overwritten. The
write itself is a bare writeFileSync on the gated path — no shell, no
exec, no command interpolation.
Every turn is sent to the engine over one of three transports: the long-lived ACP subprocess (mcode acp), the one-shot exec subprocess (mcode exec), or — added in slice S2 — an in-process runtime host (packages/webui/server/lib/runtime-host.js) that owns the same CliService the TUI does. The choice is made server-side, per turn, before the engine spawns — and it is decided in two different places, evaluated in this order:
process.env.MCODE_USE_ACP === "0"forces the exec transport (packages/webui/server/routes/chat.js#handleSend; the code comment there calls it the escape hatch for an ACP protocol regression). This is the only read of the variable in the codebase.MCODE_USE_ACP=0short-circuits all three transports.- Otherwise the turn is handed to
runMcodeAcp(packages/webui/server/lib/mcode-acp.js), which silently re-routes torunMcodeExecin its first branch whencs.permissionsis set and is anything other than"Full access". - Only when neither applies does the turn actually run over ACP.
S2 (slice 2 of the runtime-first migration) introduces a new switch alongside MCODE_USE_ACP:
| Env var | Default | Accepted values | What it does |
|---|---|---|---|
MCODE_USE_ACP |
unset | 0 → exec escape hatch (overrides everything); 1 → no effect; unset → no effect |
Today-only escape hatch; see rows below. |
MCODE_WEBUI_TRANSPORT |
acp |
acp (today's behaviour), exec (no-op — no production route consumes this value; exec transport today is reachable only via MCODE_USE_ACP=0), runtime (opt-in to the S2 in-process host; S3+ lights catalogue traffic) |
Selects the engine transport. Default keeps every response field-identical to today's main; runtime opts catalogue traffic (list/title) into the in-process host once S3 lands. |
Resolution rule, in priority order:
MCODE_USE_ACP=0⇒exec, regardless ofMCODE_WEBUI_TRANSPORT. The legacy escape hatch wins.MCODE_WEBUI_TRANSPORT=exec⇒ no-op. No production route consumes this value; theexectransport today is reachable only viaMCODE_USE_ACP=0. Documented so the contract does not drift when a future slice wires the value.MCODE_WEBUI_TRANSPORT=runtime⇒runtimefor catalogue traffic (S3+); active turns still go through ACP today (S4 wires them). Per-call fallback to ACP on any runtime-host failure so a runtime regression never breaks the sidebar.MCODE_WEBUI_TRANSPORT=acp(default) ⇒ today's ACP path. Permission-mode re-route still applies.- Unknown value (e.g. typo) ⇒ falls back to
acpwith a one-line warning to stderr. The server never refuses to boot because of an unknown transport.
| Turn condition | Transport | Decided at |
|---|---|---|
MCODE_USE_ACP=0 in the server environment |
exec | routes/chat.js#handleSend |
MCODE_WEBUI_TRANSPORT=exec |
(no-op — same as default acp; exec transport today is reachable only via MCODE_USE_ACP=0) |
server/lib/config.js#MCODE_WEBUI_TRANSPORT (no route reads this value) |
cs.permissions is Ask, Auto, or Read (not Full access) |
exec (silent re-route inside ACP entry) | mcode-acp.js#runMcodeAcp |
MCODE_WEBUI_TRANSPORT=runtime (S3+) |
runtime for catalogue traffic (sessions list/title); active turns remain on ACP until S4 | server/lib/acp-client.js#listAllMcodeSessions / #getMcodeSessionTitle (ACP fallback on any runtime-host failure) |
otherwise — factory default is permissions: "Full access" (server/lib/state-bus.js initial state) |
ACP | mcode-acp.js#runMcodeAcp |
S2 invariants (must remain true on every later slice):
- Default
MCODE_WEBUI_TRANSPORT=acpis field-identical tomain. No existing endpoint response may shift; no child process count may grow. The verification suite proves this on every commit by running the full webui node:test suite with no env override. - S2 ships the host skeleton.
createCatalogueHostandcreateTurnHostare exported fromserver/lib/runtime-host.js. S3 wires the catalogue path (list/title) intoacp-client.js; S4 wires active turns; S5 wires models; S6 wires interactions/accounts. S7 flips the default toruntime. - The catalogue host hands back the runtime's application handles, and they are process-local.
createCatalogueHostreturnsadapter,cliService,apiHost,controller,application,applications, andclose.applicationis the process-local product facade (events/models/skills/plugins/permissions/ …);applicationsis the feature-application tree. Turn diff lives on the latter only:applications.session.diffcarriesgetTurnDiff/revertTurnDiff/reapplyTurnDiff, and the process-local facade declares no diff member at all, so diff read offapplicationis alwaysundefined. Both handles stay inside the server process — no route serves them today, and a slice that turns diff into an endpoint must authorize and narrow the surface rather than forward the whole tree. - The catalogue path is opt-in via
MCODE_WEBUI_TRANSPORT=runtime. Setting it lights the catalogue host for list/title; an in-flight failure falls back to ACP for that one call so a runtime regression never breaks the sidebar. The cache (mcodeSessionsCache) is shared between paths, so a sidebar fetch served by either path serves the next read equivalently. - The catalogue projection mirrors the ACP adapter's
toAcpSessionInforule set (server/lib/catalogue-sessions.js; rules live inpackages/tui/src/acp/agent.ts, predicates inpackages/tui/src/runtime/delegation.ts): internal sub-agent sessions (workerpurposeprefixeslocal-task:/local-background-task:/team-plan:,sessionKind: "task", or a builtin sub-agentagentName—explore/worker/verifier) and sessions with a missing or non-absolutecwdnever reach the sidebar, exactly as ACP drops them; an emptytitleand a missing timestamp omit their keys —title: nullnever appears on the wire.catalogue-via-runtime.test.jspins this field-for-field against an independently re-derived expected page, and its zero-spawn probe is baseline-relative: descendants already alive when a window opens count as environment noise, while a genuine in-window spawn still trips the assertion. - R1 mitigation (process-isolation loss) lives in the turn host. Every call into
adapter.sendMessageis wrapped so a runtime-side throw becomes a stream-shaped error frame and never escapes the turn. Tests inpackages/webui/test/server/runtime-host.test.jspin this with a mutation that drops the inner catch — the test goes red if the boundary is removed. - R2 mitigation (abort semantics) lives in
createTurnHost#abortSession. It returns{success:true, elapsedMs}after at most a 5 s wait for the stream to settle; it does NOT rely on subprocess kill, because there is no subprocess. The bound keeps graceful shutdown responsive even on a wedged runtime. - R8 mitigation (wedged host) lives in
createCatalogueHost#close. It racesapiHost.close()against a 5 s timeout so a wedged dependency chain cannot wedge webui's graceful shutdown.
Contract notes:
- There is no
/execcommand. The webui button-command set isCMD_BUTTON_COMMANDS—new,clear,status,sessions,review,help,usage,stop(server/lib/interaction/command-registry.js) — and it is the only declaration of that set. The command cache inserver/lib/acp-client.jsfills itswebuigroup from it, so/helpand the composer's slash palette name exactly the commandsPOST /api/cmdaccepts,/reviewincluded. (A second, hand-written list inacp-client.jsused to omit/review, which is why/helpand the palette disagreed with the 400 branch; it is gone, andpackages/webui/test/lib/command-list-drift.check.mjspins the relationship so it cannot come back.) Transport is never switched by a slash command; the two conditions above are the whole rule. - The permission mode is selectable in the composer (Ask / Auto / Full access;
webapp/components/composer.tsx#PERMISSION_MODES) or viaPOST /api/permissions, which also acceptsread. The route writes the label intocs.permissionsunconditionally (server/routes/model.js#handleSetPermissions) — that label is what steers the next turn's transport. - An exec turn is not a degraded permission mode: the mode still reaches the engine as the
--permissionspawn flag (Ask→ask, Auto→auto, Read→read, elsefull; the mode mapping inmcode-exec.js), the session continues via--session, and the recorded model is passed via--model. - A live exec child has no RPC surface:
session/set_config_optioncalls (model, permission) returnno_acp_sessionand take effect on the next turn (server/lib/mcode-rpc.js#noLiveClientFailure); the same call lands on the live child immediately on an ACP turn. Warning semantics are documented inpackages/webui/docs/API.mdunderPOST /api/permissions.
Known costs of an exec turn — all of these are current behaviour of this tree, not planned fixes:
- No tool-call lines:
collectExecResultconsumes onlydelta/message/exec.resultstream events, so→tool rows never appear (the ACP path renders them viaapplyToolUpdate). - The thinking-effort pick is not transferred:
applyRecordedModelruns only on the ACP path, andbuildExecArgshas no thinking flag — the engine runs its own default. - No session-title write-back outside the finalize path: the title write-back runs only on the ACP finalize path (
getMcodeSessionTitleis also read on the session-switch fallback and the on-demand title API, but only finalize persists). - No interactive channel: the child's stdin is closed immediately after the prompt is written, so engine-side questions cannot reach the browser; questionnaire-type turn errors surface as alerts with a hint to re-ask via the composer (
routes/chat.js#handleSenderror branch).
This section records what the current source tree does, not a frozen contract. During a turn the two transports are distinguishable in the process list: an mcode … acp child is an ACP turn, an mcode … exec --input - child is exec. The operator-facing view — when you hit each transport, what it costs, and what to do — is the transport section of webui.zh-CN.md.
The server carries an internal engine layer, packages/webui/server/engine/, whose first job is a capability declaration: every engine surface webui is wired to declares, as a reviewed module constant, which of 14 capability keys it supports and — for partial support — exactly which sub-items are missing. The design decision and the audited matrix behind every value live in doc/engine-abstraction-design.md (out-of-tree working document); the declaration itself is the code source of truth.
Why declarations instead of try-and-see: a missing capability must be a fact the UI can read before calling, not an exception discovered mid-call, and it must never be a silent empty implementation — an empty list or {ok:true} would tell the user "succeeded with nothing", the fake-success failure mode fixed in #110 and refused here by construction.
The 14 keys (one per row of the design matrix; key ↔ matrix row in parentheses):
sessionCrud (会话 CRUD), streamingSend (流式发送), interrupt (中断), toolSkillInvocation (工具/技能调用), turnDiff (回合级 diff 查询), turnRewindRedo (回合撤销/重做), plugins (插件管理), mcp (MCP), subagents (子 agent), usageStats (用量统计), authCredentials (认证/凭据), updateCheck (更新检查), fileReadWrite (文件读写), gitOperations (Git 操作).
Levels and rules (server/engine/capabilities.js):
full— the surface is complete.partial— must enumeratemissingsub-items and carry areason. Never "half works, nobody knows which half".none— must carry areasondistinguishinginterface-absent(no such method on the surface at all) fromimplementation-absent(the layer above has it, this surface does not open it).servedBy— optional, and only on anoneentry. Names the provider whose in-process host actually answers the request when this provider does not implement the capability itself. Rejected onfullandpartial(a provider that partly implements a capability is not "served elsewhere"), and aservedBynaming an unregistered provider is a boot-time throw, not a runtime 404.
Current declarations (the two runtime surfaces transcribed from the audited matrix and re-verified against the live method surfaces at the 26043e9b baseline — 91 adapter methods, 94 CliService methods plus the applications.session.diff facade):
| Key | local-runtime-v2 | tui-runtime-adapter |
|---|---|---|
| sessionCrud | full | full |
| streamingSend | full | full |
| interrupt | full | full |
| toolSkillInvocation | partial — missing setMode (no session-mode write; M3-B9) |
partial — missing setMode |
| turnDiff | full | none (implementation-absent on the adapter) |
| turnRewindRedo | full | partial — missing reapplyTurnDiff |
| plugins | full | partial — missing previewGithubPlugin, importGithubPlugin, listEnabledPlugins |
| mcp | full | full |
| subagents | partial — missing getDelegationSnapshot, stopDelegation (they live on the adapter's access-context, not the CliService surface) |
full |
| usageStats | full | full |
| authCredentials | partial — missing setConfigOption (the GENERIC config write; M3-B9) |
partial — missing setConfigOption |
| updateCheck | none (interface-absent) | none (implementation-absent) |
| fileReadWrite | partial — missing file-write |
partial — missing file-write |
| gitOperations | partial — missing git-diff, git-commit, git-branch |
partial — missing git-diff, git-commit, git-branch |
The third registered provider is the acp transport (M4-1) — a transport rather than an in-process surface, declared in server/engine/providers/acp.capabilities.js and audited against MCODE_ACP_CAPABILITIES, the protocol's live wire table, because a subprocess has no object to reflect:
| Key | acp |
|---|---|
| sessionCrud | partial — missing deleteSession, renameSession, archiveSession (session/new · load · list · close · resume · fork · activate are on the wire, and session/delete is registered with no handler) |
| streamingSend | full (session/prompt) |
| interrupt | none — interface-absent: session/cancel IS registered, but it is a notification, and a delivered cancel certifies that it was sent, never that the turn stopped |
| toolSkillInvocation | partial — missing listSkills, listRuntimeSkills (the protocol has no skill enumeration) |
| turnDiff | none — interface-absent, servedBy local-runtime-v2 |
| turnRewindRedo | none — interface-absent |
| plugins | none — interface-absent, servedBy local-runtime-v2 |
| mcp | partial — missing mcp-configure, mcp-inspect, mcp-clear, mcp-list (MCP servers take effect inside a turn; nothing configures or inspects them) |
| subagents | partial — missing getDelegationSnapshot, stopDelegation, listBackgroundTasks (activity is parsed off the event stream only) |
| usageStats | partial — missing getSessionUsage, getSessionUsageSummary, watchSessionUsageCommits (plan quota is queryable over mcode/account/status; the token detail webui shows beside it is read from the runtime DB, not the engine) |
| authCredentials | partial — missing the OAuth flow, the API-key surface and the user model-provider CRUD. The config-option write IS present, and dispatches the model and permissionMode config ids |
| updateCheck | none — interface-absent (available_commands_update refreshes the advertised command catalogue, which is not an update check) |
| fileReadWrite | none — interface-absent (webui's /api/fs family is its own node:fs implementation) |
| gitOperations | none — interface-absent (webui's /api/git family wraps the OS git binary) |
Two cells are stronger here than on either runtime surface, and flattening them would be the unearned claim the design matrix forbids: the protocol registers session/set_mode as a real request, so toolSkillInvocation does not miss setMode over acp; and session/set_config_option dispatches the model and permissionMode config ids, so two of the three bridged writers of MODE_WRITE_BRIDGED_CONFIG_IDS are genuinely reachable.
The fourth registered provider is the exec transport (M4-2) — the one-shot mcode exec subprocess, the third legal MCODE_WEBUI_TRANSPORT value. It is not a mode of acp and not an alias for the tui package: mcode-exec.js writes the prompt to stdin and parses stream-json off stdout, so there is no request channel and therefore no methods to call. What it has instead is CLI options and event types, and that is what server/engine/providers/exec.capabilities.js records in EXEC_INTERFACE and audits against:
| Key | exec |
|---|---|
| sessionCrud | partial — missing createSession, listSessions, getSession, updateSession, renameSession, archiveSession, deleteSession, forkSession, getSessionForkOptions, loadSession, activateSession. Only --session / --continue exist, and they re-enter a session rather than choose one |
| streamingSend | full (--input - plus the stream-json event stream) |
| interrupt | none — interface-absent: nothing to call. The SIGINT/SIGTERM/SIGHUP in packages/tui/src/cli/run-exec-command.ts are signals webui sends to the child it spawned, i.e. webui's own kill cascade, not a capability the transport offers |
| toolSkillInvocation | partial — missing listSkills, listRuntimeSkills, listPendingPermissions, replyPermission, setMode. The transport PRODUCES tool_call items and webui consumes but does not render them; and --permission is fixed at spawn time (the CLI says ask requires TUI/ACP), so there is no permission request/reply pair |
| turnDiff | none — interface-absent, servedBy local-runtime-v2 |
| turnRewindRedo | none — interface-absent (mcode exec review reviews local git changes and carries no turn coordinate, so it is not a rewind surface) |
| plugins | none — interface-absent, servedBy local-runtime-v2 |
| mcp | partial — missing mcp-configure, mcp-inspect, mcp-clear, mcp-list (--config can hand the process an MCP configuration; nothing configures or inspects one afterwards) |
| subagents | none — interface-absent: the event union has no delegation kind, and packages/tui/src/headless/runner.ts refuses to open sub-agent Sessions at all |
| usageStats | partial — missing getSessionUsage, getSessionUsageSummary, watchSessionUsageCommits. Stronger than acp here: turn.completed.usage is emitted on the wire, so this transport has something under its three missing names and acp does not |
| authCredentials | none — interface-absent: with no request channel there is no mcode/account/status and no session/set_config_option. --model / --effort are per-run spawn flags, not a readable or writable account surface |
| updateCheck | none — interface-absent (mcode update is a sibling CLI command, and there is no channel to notify over) |
| fileReadWrite | none — interface-absent (--file attaches a file to the prompt; webui's /api/fs family is its own node:fs implementation) |
| gitOperations | none — interface-absent (webui's /api/git family wraps the OS git binary) |
Three cells are weaker than acp, and for structural reasons rather than unfinished engine work: interrupt (acp has a cancel notification, exec has nothing to declare one on), subagents (acp can parse sub-agent activity off its stream, exec's event union has no such kind) and authCredentials (acp has two RPC methods, exec has no channel). The reverse exception is the same two keys for the same reason, which is a finding rather than a copy: /api/turn-diff* and /api/plugins* project the in-process v2 host and gate on no transport, so every transport inherits it.
The mismatch this audit surfaced, and what it cost. The three event names collectExecResult branched on — delta, message, exec.result — are the supervisor's internal stream-event names. --output-format stream-json writes only what ExecEventProjector produces (packages/tui/src/headless/output.ts refuses the format with no projector, packages/tui/src/headless/runner.ts always supplies one, and the encoder's result() leg goes through projector.complete() rather than writing the ExecResult itself), so the wire carries the ten ExecEvent types and the two name families did not intersect.
That was not a missing feature but a dead data plane, and it is worth spelling out what a user saw on MCODE_USE_ACP=0 — and on every non-Full access permission mode, which silently re-routes to exec: reasoning and answer text streamed in and never landed, the turn ended with no answer line at all, the context counters never moved, and the follow-up turn started a brand-new engine session because the session id was never read back. D1 rewrote the consumer against the wire:
| Wire event | Consumed as |
|---|---|
sessionId on every line |
the engine session this run entered, written back to cs.mcodeSessionId so the next turn continues it |
item.started / item.updated |
item.contentDelta appended to the answer / reasoning accumulator and streamed as a ● / ▲ line |
item.completed |
item.content, adopted only for an item that never streamed a delta |
turn.completed |
per-turn usage and durationMs |
turn.failed |
status and error |
exec.completed |
the terminal ExecResult, and the call to finalize() |
EXEC_INTERFACE.consumedEvents now names exactly those six, EXEC_INTERFACE.baseOnlyEvents names the four that carry nothing beyond ExecEventBase, and a test asserts the two partition the union, that the parser's switch arms are the same set, and that every consumed name is one the wire can emit. A tool_call item is consumed but not rendered — it carries a toolCall payload rather than text — which is why toolSkillInvocation stays partial.
servedBy is the plan's one reverse exception, and it is load-bearing. turnDiff and plugins are honestly none on the protocol, and the three /api/turn-diff and ten /api/plugins endpoints still work on the default acp transport, because they project the in-process local-runtime-v2 host through getEngineCatalogueHost() and are gated on no provider declaration. Reading the level alone would eventually 501 two working features the moment a frontend consulted the transport's provider instead of the default one. summarizeCapabilityHosting(capabilities) and resolveCapabilityHostProvider(providerId, key) expose the routing fact; the hosted keys deliberately stay in summarizeUnavailableCapabilities, because the provider really has none and that {none, partial} shape is already on the wire. The exec provider carries the same two fields for the same structural reason, which is what makes servedBy a per-key declaration field rather than an acp special case.
Registering a provider is not routing to it. Every capability gate resolves its provider through a transport→provider table in its own family module, and none of them lists acp or exec: a miss there means "no provider claims this transport yet", and the gate passes. So neither M4-1 nor M4-2 changed any gate's verdict on any transport. Making a transport provider actually reachable — chat.js transport selection reading the registry — is M4-3, and the test that keeps the two apart walks all sixteen resolve*Provider functions on both unwired transports.
Auditing a table that cannot be imported. The acp declaration is checked against MCODE_ACP_CAPABILITIES, a live constant the routes read. The exec contract is TypeScript in another package, and importing it would put @mavis/* on the boot path, so EXEC_INTERFACE is transcribed — and transcribed tables rot. Two live cross-checks keep it honest: every option in applyExecCliContract, every type in the ExecEvent union, every ExecItem kind and every signal run-exec-command registers are read out of the real sources and compared, and buildExecArgs() (a pure function) is invoked so the table can never shrink below what webui actually sends. auditExecCapabilities itself has two rules rather than three, and the missing third is a decision: "a partial's missing must not name a mechanism the interface exposes" is vacuous here, because missing names provider methods while the coverage table names interface mechanisms, and the two namespaces cannot intersect. A check that cannot fail reads as coverage in a file whose only job is honesty.
Read-only, declaration-backed (boots no host, runs no probe). Returns one provider's declaration plus the degradation summary the future capability-driven UI renders from:
GET /api/engine-capabilities[?provider=<id>]
200 { ok, provider, transport, capabilities: { <key>: {level, missing?, reason?} × 14 },
unavailable: { none: [key…], partial: [{key, missing}…] } }
404 { ok: false, code: "unknown_engine_provider", knownProviders: [...] } // caller confusion
Default provider is local-runtime-v2 — unchanged since B1, and deliberately so: M4-1 and M4-2 each added a provider, not a default, so every existing caller (including the webapp's own degradation test) keeps seeing the declaration it saw before. The registered ids are local-runtime-v2, tui-runtime-adapter, acp and exec. Unknown ?provider= answers 404 with the id list — it cannot collide with the 501 reserved for engine limitations.
server/engine/errors.js defines EngineCapabilityNotSupportedError (structured: capability, provider, missing, reason). assertEngineCapability(capabilities, key, provider, subItem?) throws it for level none and for the missing half of a partial. Both HTTP layers (server/app.js#invokeHandler and the legacy server/router.js dispatcher, same centralisation as the existing 413 body-cap mapping) turn it into:
501 { ok: false, code: "engine_capability_not_supported", capability, provider, missing?, reason?, error }
501, not 400/404/500: the request was well-formed; the engine provider lacks the feature. This mirrors the existing unsupported → 501 mapping in routes/protocol.js. The frontend treats engine_capability_not_supported as expected degradation (hide the entry point per the level table), never as an error toast.
POST /api/protocol/set-mode (#67) and POST /api/protocol/set-config-option (#68) sit behind a hard capability gate, and they are the first endpoints in the migration whose answers change for some deployments. The change has exactly one trigger — the connected engine provider declares the capability absent — and it is worth being precise about, because everything outside it is unchanged byte for byte.
| Request | Before | After |
|---|---|---|
#67, provider declares toolSkillInvocation.setMode |
forwarded to the engine; whatever it answered | 501 {ok:false, code:"engine_capability_not_supported", capability:"toolSkillInvocation", provider, missing:["setMode"], reason, error} |
#68 with any config id other than model / permissionMode, provider declares authCredentials.setConfigOption absent |
forwarded to the engine; whatever it answered | 501 {… capability:"authCredentials", missing:["setConfigOption"] …} |
#68 with model or permissionMode |
forwarded to the engine | unchanged — the bridge below |
any of the above, the engine itself answers unsupported |
501 {ok:false, code:"unsupported", fallback:"send_plan_as_prompt"} |
unchanged, including the fallback field |
any of the above, the provider does not declare the capability absent (including every request on the default acp transport) |
unchanged | unchanged |
Two consequences of that table are deliberate rather than incidental:
- The capability 501 carries no
fallback. The hint is the degraded action for a feature that exists and whose call failed. Where the engine has no mode write at all there is nothing to degrade to, and advertisingsend_plan_as_promptfrom a "this is not available" response would offer a workaround for a missing feature. The engine's ownunsupportedrefusal keeps its hint. - On the default
acptransport nothing changes at all. No provider is registered foracpuntil migration step M4, so the gate reportsunregistered-transportand every response is the pre-M3 one. The refusals above are reachable on theruntimetransport, wherelocal-runtime-v2is the registered provider.
The bridge. A provider can refuse the generic config-option write and still have the dedicated writers webui's own controls depend on. #68's gate therefore asks for a sub-item derived from the request: model asks for selectModel, permissionMode asks for setPermissionMode and thinkingEffort asks for setThinkingEffort, all of which pass a provider that denies setConfigOption; every other config id asks for setConfigOption and gets the 501. The exemption is exactly three named ids — never a prefix, never a default — and it does not survive a none: a provider with no authCredentials at all has no dedicated writer either. (The third id arrived in M3-B14; see below.)
What the user sees. The permission-mode selector and the model selector are hidden, not disabled and not accompanied by an error message (webapp/lib/engine-capabilities.ts, wired in webapp/components/composer.tsx). A toast would report a failure for something the user was never able to do, offer nothing to act on, and reappear on every click. The rule is fail-open: the controls are shown until the declaration positively says the engine cannot do it, so a failed or slow /api/engine-capabilities request never removes a working control.
POST /api/set-model (#58) and POST /api/permissions (#59) are the second half of the model endpoint family; B4 moved its read, this batch moves the write. The reasoning leaves the route and lands in packages/webui/server/engine/model-writes.js, where it is named, exported and tested on its inputs.
Nothing a client can observe changed. Every status, response field, field order, warning string and engine push — including which push happens first and what it is allowed to say when it fails — is the one these two endpoints produced before. The boundary is:
| Concern | Home after B10 |
|---|---|
| webui id → engine wire value | resolveEngineModelConfigValue |
| variant channel vs effort channel, and the push order each implies | planModelSelectionPush |
the set_config_option calls |
pushEngineModelSelection / pushEnginePermissionMode |
| permission mode → label / engine value | resolvePermissionSelection |
the rule for when the local configOptions snapshot may claim the engine's new effort |
applyThinkingEffortMirror (the write stays in the route — cs is webui's own state) |
body parsing, the 400s, the cs.model / cs.permissions writes, pushStateFor, the response bodies |
packages/webui/server/routes/model.js |
Two forms the picker deals with are deliberately different and stay that way. What the engine receives is the wire form — m:<provider>:<model>:u, or m:<provider>:<model>:v:<variant> for a switchable builtin, plus a bare level for thinkingEffort and an engine vocabulary word for permissionMode. What webui records is the user-facing form — cs.model.name in <providerKey>/<engineModelKey>, cs.model.thinking, cs.permissions as a label. The map between them is what the suite pins, field by field, over one row per (engine option shape × request shape) in packages/webui/test/lib/engine/model-writes.test.js.
The variant channel (ticket 36) is unchanged and now covered by name. A switchable builtin (thinking_config.mode: switchable, e.g. MiniMax-M3) has no engine effort vocabulary — the engine rejects every thinkingEffort value for it — and advertises it only as the wire pair v:thinking / v:none-thinking. Such a pick is therefore one model push carrying both the model and the on/off level, with no second push at all. Every other model keeps the two-push contract: model first, then thinkingEffort, because the engine rejects an effort set when no model is selected. A cleared level on the variant channel means the engine's default variant, not "off" — the normaliser only knows on and off.
The 4-second SSE race window is unchanged, and now has both halves pinned. A pick stamps the fields the request actually carried, all with one timestamp, so applyConfigOptionUpdate's ownership-aware mirror (server/lib/mcode-acp.js, ticket 08) defers the engine's wire-form echo for 4 seconds instead of letting it overwrite the chip a few milliseconds after the optimistic write. A field the request did not carry is not stamped, so a later cross-client change to that field still mirrors immediately.
contextWindow is still recorded and never pushed. The engine's ACP surface has no channel for it, so the pick is a webui-side preference the picker reflects immediately.
These two endpoints were not gated in this batch, and that was an open decision rather than an oversight. #59 writes permissionMode only, so gating it on authCredentials.setPermissionMode would be behaviourally inert today and safe against the shipped UI (the permission selector is already hidden under exactly that declaration) — it is one assertEngineCapability call. #58 also writes thinkingEffort, which was a generic config id: gating it the same way would make the thinking-effort control answer 501 for the same reason #68 does for an unrecognised id. Both branches were costed in the KNOWN DEBT section of model-writes.js — bridge thinkingEffort as a third bridged id, or accept the 501 and extend the frontend's degradation to a third control. M3-B14 took the first branch, and the gate landed with it; #58 keeps its pre-B10 behaviour only on the paths that never reach the engine.
The bridge is no longer an unverified exemption. selectModel and setPermissionMode — the first two sub-items MODE_WRITE_BRIDGED_CONFIG_IDS named — are in the snapshot audit's REQUIRED_METHODS, so a real booted host is checked for both of them on the adapter and the CliService surface, and a declaration that stops listing one goes red. Neither surface carries a setThinkingEffort / selectThinkingEffort, which is the fact the gating decision above turns on. The third id M3-B14 added points at that same absent method, so the audit tracks it as a proven absence rather than as a presence — see the M3-B14 section for what that means when the engine ships the writer.
M3-B11: the provider family moves behind the facade, and the two provider files become one (storage change)
GET /api/providers (#62), PUT /api/providers (#63), POST /api/providers/test (#64), GET /api/providers/presets (#65) and POST /api/providers/preset/:id/enable (#66) are the last catalogue family in the migration, and the only one that changes where a user's data lives.
What changed. webui kept two files describing the same providers: ~/.mcode-webui/providers.json (the v2 catalogue, ordered, lossless) and the engine's <engine data dir>/config.yaml custom_provider tree (a projection of the first, written by a double-write that had no transaction across it). The projection was lossy and the loss was invisible precisely because nothing read it back: a disabled provider, a coding-plan provider, a preset name and the gemini-vs-openai protocol distinction all vanished on the way to the engine, and the catalogue's ordering came from the file that was about to stop being authoritative. There is now one file. Each webui-managed entry carries its webui record beside its engine fields:
custom_provider:
acme-gateway:
name: Acme Gateway
kind: custom
api: openai-completions
options: { apiKey: …, baseURL: …, authMode: api-key }
models: { glm-5.3: { limit: { context: 128000 } } }
_webui_owned: true # ownership: webui wrote this entry
_webui_provider: { … } # the authoritative v2 record, verbatimBoth marker fields are ignored by the engine, which parses config.yaml through js-yaml with no schema rejection and reads named fields. A provider the engine cannot express still gets its key, its marker and its record — it simply has no engine fields, which is the whole difference from the double write.
The migration, and the fallback. While the store carries no _webui_provider_migration marker, the deprecated providers.json is still the authority; webui folds it into the store on the next read and stamps the marker on success, after which the file is never read again. A migration that fails — an unparseable config.yaml, a write that could not complete — leaves the store byte-identical and the old format readable, and the next read retries. The marker is a field rather than an inference ("the tree has webui entries") for one concrete reason: an operator who deletes every provider leaves a tree with no webui entries, and an inferred marker would hand authority back to the stale file and resurrect what they had just removed.
Field-by-field equivalence and both fallback paths are pinned in packages/webui/test/lib/engine/provider-migration.test.js, on a fixture built to break every assumption the migration could be quietly making: several providers, every schema field, and the boundary values (empty label, absent preset, disabled, coding-plan, the gemini protocol, a zero context limit, empty thinking levels, a model id the engine key grammar rejects, unicode, a 4096-character key).
PUT atomicity is now structural. There is one file and one rename, so the two-file disagreement the old arrangement allowed — the catalogue committed, the engine projection failed, a 200 with a warning nobody had to read — cannot be constructed. A refused write (an unparseable config.yaml is refused, never overwritten, because rewriting it would destroy every engine setting the store does not own) or a failed write leaves the previous document intact, and a concurrent reader always sees a whole catalogue.
The gates. The two write endpoints declare authCredentials and gate hard on updateUserModelProvider / createUserModelProvider: the catalogue the operator is about to see is read by the engine, so a provider that cannot write providers cannot truthfully answer 200. The three read endpoints declare the same capability and gate soft — a provider with no provider surface still serves a well-defined catalogue, so hard-gating them would delete a working UI over an enrichment. As in B9, an unregistered transport (acp, until M4) is not a 501.
What a client observes. The endpoint shapes, statuses, masking rule, keep-key convention, probe semantics and the providers.updated SSE frame are unchanged. Two response values moved with the storage: PUT's path is now the engine's config.yaml, and it also reports engineSync: {ok, written, keys} for the store write itself. GET's sources and userPath are unchanged in both field and value — they still name the deprecated file, because "which files did the server resolve" is a question an operator asks when a provider is missing, and the answer is now carried by the bilingual docs rather than by a renamed field.
Three decisions are recorded rather than taken. POST /api/providers/test names testUserModelProvider in its gate, and that method cannot answer it: the engine's tester is keyed on a persisted provider, while the endpoint tests an unsaved candidate from a form. The probe stays webui-local, which is also the only option that keeps its two load-bearing properties (the local key-format check runs before any network call, and the apiKey goes to the configured baseURL and nowhere else). The preset gallery is still webui's own template list, and the engine has a different one; the two are not the same taxonomy, so the plan's "align the two template sets" is made visible rather than closed. And a webui provider whose engine key collides with an operator's hand-written entry still overwrites it, because the key is the runtime id and a silent rename would turn a recorded model pick into an unresolvable one. All three are costed in the KNOWN DEBT sections of provider-reads.js and provider-writes.js.
M3-B10 moved these two endpoints behind the facade and left one decision open. This batch closes it, and the part worth reading is why the obvious gate on #58 would have been wrong.
The decision that was open. #59 writes permissionMode and nothing else, so gating it on authCredentials.setPermissionMode is one call and no behaviour change. #58 also writes thinkingEffort, and that config id was generic — the one the plan (§3a, row 68) says has nowhere to be delivered under a provider with no generic write. Gating #58 the same way would have made the thinking-effort control answer 501 for exactly the reason #68 does for an unrecognised id. Two branches were costed: bridge thinkingEffort as a third bridged id, or accept the 501 and hide the control. The bridge was chosen, and the exemption list is now three names:
| config id | #68 asks for | #58 asks for | what the engine receives |
|---|---|---|---|
model |
selectModel |
— the model push rides the model capability | m:<provider>:<model>:u, or :v:<variant> for a switchable builtin |
permissionMode |
setPermissionMode |
— #59 is its own endpoint | an engine vocabulary word |
thinkingEffort |
setThinkingEffort |
setThinkingEffort, on the effort channel only |
a bare level |
| anything else | setConfigOption → 501 |
— | — |
The exemption is still three named ids — never a prefix, never a default — and it still does not survive a none.
Why #58's gate is on the effort channel and not on the endpoint. #58 has two channels, and they use different capabilities. A switchable builtin (ticket 36) has no engine effort vocabulary at all, so one model push carries the model and the on/off level. The level there rides the model capability, and gating it on an effort sub-item would 501 a model switch for a capability the switch never uses. A model-only pick on the effort channel has no effort write to gate either. The predicate is therefore read off the plan rather than off the request's fields — Boolean(plan.thinkingPush) — and the variant channel is ungated by construction rather than by a second condition somebody has to keep in sync:
flowchart TD
A["POST /api/set-model"] --> B{"a live session?"}
B -- no --> B1["200 + local-only warning<br/>nothing reaches the engine, so there is<br/>nothing for a gate to be honest about"]
B -- yes --> C{"variant channel?<br/>(switchable builtin)"}
C -- yes --> D["one model push carries<br/>model + on/off level"]
C -- no --> E{"plan.thinkingPush<br/>non-null?"}
E -- no --> F["model-only or cleared effort<br/>NOT gated"]
E -- yes --> G{"authCredentials .<br/>setThinkingEffort"}
G -- allowed --> H["model push, then<br/>thinkingEffort push"]
G -- denied --> I["501 engine_capability_<br/>not_supported"]
A pure model switch answering 200 while an effort write on the same session, the same provider and the same request frame answers 501 is not an inconsistency — it is the point, and both directions are pinned in packages/webui/test/lib/engine/model-writes.test.js.
What a user sees. One change, and it is a UI change rather than a status change: the thinking-effort selector is now the third control the engine-capability rule governs (webapp/lib/engine-capabilities.ts, wired in webapp/components/composer.tsx). Under a provider that declares the dedicated effort writer absent it is hidden, not disabled and not accompanied by a message, for the same reason the other two are. No registered provider declares it today, so no control disappears on the current builds; the rule stays fail-open, and a failed or slow /api/engine-capabilities request still shows everything.
What changed for #68. POST /api/protocol/set-config-option with key: "thinkingEffort" no longer answers 501 under such a provider. Nothing in the shipped webapp calls #68, so there is no client to break, and the change makes the two endpoints agree: a config id must not be deliverable through #58 and refused through #68 for the same provider. contextWindow is the honest generic example now, and both the suite and this table say so.
The name is a forward contract, and the audit says so rather than implying otherwise. selectModel and setPermissionMode are methods the audited surfaces really carry, which is why B10 could add them to the snapshot's REQUIRED_METHODS and have the audit check them on the adapter and the CliService. setThinkingEffort is not. The snapshot test probes the real booted host by reflection and asserts its absence on both surfaces, in a new unimplemented list that means precisely one thing — this surface must not carry this method — and that turns the audit red the moment either surface grows one. That is the whole closure mechanism, and it is deliberately one-directional: the engine shipping a dedicated effort writer is an event nobody here can schedule, and the audit is what makes it impossible to miss. When it happens, the name moves from unimplemented to methods, the declaration is re-audited, and the control comes back on its own.
What is deliberately not done: no provider's authCredentials declaration was edited to list setThinkingEffort in missing. Listing it would make every provider refuse the effort write and remove the control for every user today — the other branch's cost, not this one's. The gate reads the declaration, the declaration describes the surface, the surface really has no such method, and the gate is therefore inert. That is the truthful state of the world rather than a faked one.
- M1 done in this batch: host construction (
createCatalogueHost) moved verbatim intoserver/engine/providers/local-runtime-v2.js;runtime-host.jsre-exports it, so every existing importer is untouched. No existing route's behaviour changed;GET /api/engine-capabilitiesis a new, additive endpoint. - M2 done (declaration-vs-implementation snapshot):
packages/webui/test/lib/engine/capability-snapshot.test.jsboots a REAL catalogue host on an isolated tmp data dir (MINIMAX_DATA_DIR+ everyMCODE_WEBUI_*path pinned before the provider import) and audits everyfull/partialkey of both providers —fullrequires every tracked method to exist on the declared surface (adapter/cliService/applications.session.diff),partialrequires the present half to exist, the method-namedmissingitems to be genuinely absent, and kebab-case sub-capabilities (file-write,git-diff) to have no covering method;noneis not method-checked. The tracked-method table was reflected off the live surfaces (91 adapter / 94 CliService methods), not copied from the design matrix; mutation tests in the same file pin that flipping a level, deleting a method, or growing a sub-capability each goes red. A registry-driven guard (engine/index.js#listEngineProviderIds) rejects any provider declaration carrying keys outside the 14-key contract, so a typo cannot pass silently. - Capability probing (design §2.3 step 2) is deliberately not in this batch: no route consumes a probe result yet, and wiring one would touch the catalogue host lifecycle that M1 leaves alone. It lands with the first A-batch route that needs it.
- New-provider admission rules (enforced by the snapshot tests in
packages/webui/test/lib/engine/capabilities.test.js): all 14 keys declared;partialenumeratesmissing+reason; declaration levels are pinned — a level flip without re-auditing the surface goes red in CI; calling an undeclared capability answers the structured 501, never an empty implementation.
The ACP handshake is bidirectional, and both directions are decided by one initialize payload. This section records what the webui advertises, why the list is that short, and what happens to a request the engine sends when the webui has no surface to answer it on.
packages/webui/acp.mjs sends its capabilities under clientCapabilities — the ACP v1 InitializeRequest field the engine reads (packages/tui/src/acp/agent.ts:434). The value is the exported CLIENT_CAPABILITIES constant, and today it is exactly one entry:
| Advertised | Engine behaviour it switches on | Does the webui consume it? |
|---|---|---|
plan: {} |
the plan_update session update (agent.ts:1328 gates, agent.ts:1356 sends) |
yes — streamAcpPrompt writes cs.plan (server/lib/mcode-acp.js:1138) and the plan modal renders it |
elicitation.form |
the elicitation/create request path (acp/interactions.ts:607) |
no — no form UI exists |
auth.terminal |
authMethods in the initialize response (agent.ts:455) |
no — no terminal to run mcode login in |
_meta['minimax-code/extensions'] |
goal / queue / delegation / current-session notifications (acp/extensions.ts:277) |
no — nothing subscribes to those method names |
A capability is a promise to answer, so the list carries only what the webui really consumes. The three omitted entries are not free. elicitation.form makes the engine send an elicitation/create request that this client can only decline, and the engine then dismisses the Runtime questionnaire outright (acp/interactions.ts:647) — a questionnaire the user could have answered in the TUI simply disappears. The extension _meta is pure cost with no consumer: those notifications arrive as top-level ACP methods, and the only goal_update the webui handles is a session/update sub-kind (acp.mjs:305), a different channel.
Reading a capability off the wrong field is silent, not loud. The engine's params.clientCapabilities ?? {} means a payload sent under any other key negotiates nothing, and every capability-gated projection stays switched off with no error anywhere. That was the state of this tree: the webui sent capabilities: { mcpCapabilities: … } — a key that is not a field of the ACP v1 ClientCapabilities type, holding a member the type does not declare — and the plan projection never ran.
The engine issues its own requests over the same pipe: session/request_permission, elicitation/create, fs/read_text_file, fs/write_text_file, terminal/*. McodeAcpClient#_dispatch answers every one of them. A message carrying an id, a method, and neither a result nor an error is a request. JSON-RPC ids are per-direction, so the engine's request may reuse an id the webui already used for its own outbound call, and the two spaces must not be confused.
With no clientRequest handler installed — today's state — the answer is a JSON-RPC error, -32601, naming the method. Silence is not the safe default here:
- The engine awaits these requests with only a cancellation signal (
acp/interactions.ts:562). An unanswered request occupies an interaction-scheduler slot for the life of the connection, and when the pending queue overflows the whole ACP connection is closed (interactions.ts:242,MAX_PENDING_INTERACTIONS). The transport dies rather than degrading. - A declined request is the engine's own outcome, not a new one. A request that throws resolves to
decision = 'deny'(interactions.ts:581), and a questionnaire that cannot be answered is dismissed fail-closed (interactions.ts:647).
An error rather than a synthetic "cancelled" result says plainly that this client never considered the question, and carries the method name into the engine's log. Every declined request also logs [acp] declined unhandled client request: <method> on the webui side, so an engine asking for something this client cannot do is visible rather than inferred.
The seam for a real surface is the clientRequest constructor option: (method, params) => result | Promise<result>. Its resolved value becomes the JSON-RPC result; a throw or rejection becomes an error response carrying the thrown message and, when it has one, its code (otherwise -32603). Nothing in the webui installs a handler yet — routing a decision through to the browser is separate work, and the honest current state is that the webui has no interactive surface to offer.
The engine announces its own failures on stderr and then dies; the crash alert is raised by the webui, not by the engine. McodeAcpClient therefore keeps a bounded tail of that stream — the last 2KB and the last 20 lines, cleared at every start() so one process's crash text can never be blamed on the next — and the [mcode-acp.start] and [mcode-acp.stream] error alerts carry it as data.stderrTail, prefixed with [acp stderr truncated, showing the tail] when anything was dropped. An exit code is not a diagnosis: mcode acp exited (code=1) cannot separate a lock the engine could not take from a configuration it refused to parse, while the engine's own line (agent_name_conflict_migration_failed:lock) says which.
stderrTail is additive and optional. A silent engine leaves data byte-identical to what it was before the field existed, so no consumer of the alert contract has to learn a new required key. The debug constructor option keeps its old job — mirroring the stream live to the server's own stderr as it arrives — but all three construction sites in the shipped server pass debug: false, so in a running webui the alert's tail is the only channel that stderr has.
plan: {} turns on a notification, not a question. A plan review carries a single approve option and the Runtime pins allowOther: true on every step, so the engine settles it fail-closed through the questionnaire path rather than turning it into a permission request — which is why advertising plan is safe for a client that cannot answer anything. The permission-request path is a separate switch the webui never turns on.
Two consequences are not in scope here, and both matter to whoever picks this up next:
- Receiving
plan_updateis not the same as being able to act on it. The webui's plan modal has no reachable decision channel to the engine, and the payload mapping is its own question: the engine nests the body asupdate.plan = { type, planId, content }and puts nothing at the update's top level. - The questionnaire and permission surfaces stay dark. The webui declines what it cannot answer, which is stable and honest, but it is not the same as being able to answer.
The composer mounts a thinking control only for models whose /api/models entry carries thinkingLevels. A model with no list never shows one — by design, a no-op control is worse than none. Where the list comes from, and what a pick actually does on the wire, depends on which of the engine's two thinking schemas describes the model:
| Model kind | thinkingLevels |
A pick travels as |
|---|---|---|
Provider catalogue with thinking.effortOptions (custom providers, MiniMax-M3.1-Flash-Preview) |
the engine's effort list, verbatim (e.g. default/low/medium/high/xhigh/max); thinking itself cannot be turned off (engine marks the model forced_on), only the depth is pickable |
session/set_config_option{configId:"thinkingEffort"} after the model selection |
Switchable builtin (thinking_config.mode: switchable + on/off variants, today MiniMax-M3) |
["off","on"] — a two-state toggle, never a depth scale |
one set_config_option{configId:"model"} whose value carries the variant (m:minimax_api:MiniMax-M3:v:none-thinking / :v:thinking) |
forced_on with no effort dimension (MiniMax-M2.7, MiniMax-M2.7-highspeed) |
absent — no control | n/a |
Why two channels: the engine's thinkingEffort option only accepts values the selected model advertises as effortOptions. A switchable builtin has none — the engine rejects every effort value for it (Thinking effort is not advertised for the selected model). Its on/off state is the model variant dimension, so the webui folds the level into the model selection. variantChannelFor (server/lib/engine-catalogue.js) derives the level→variant map from the engine's own variant tree (which variant disables thinking), never from hard-coded names.
Source of the builtin metadata: the engine materialises its builtin catalogue into <engine data dir>/config.yaml under provider.minimax.models (with thinking_config, variants, thinking.effortOptions). GET /api/models reads that tree on every request (readEngineBuiltinThinking) and annotates both the builtin shell entries and the engine-session wire-form entries — the latter because applyConfigOptionUpdate mirrors the engine's wire-form currentValue into cs.model.name outside the pick window, and the composer matches the active model by id.
Contract details:
POST /api/set-model{model, thinking?}recordsthinkingincs.model.thinkingwhatever the channel;thinkingSyncedreports the pick actually reaching the engine — for the variant channel it is the model push carrying the level, andmcodeSynced/thinkingSynceddescribe that one push from both angles.- Session boot replays the pick (
applyRecordedModel): effort models push model-then-effort; variant models push one variant-carrying model selection and skip the effort push. The order is load-bearing — the engine rejects athinkingEffortset while no model is selected (Select a Session model before changing thinking effort., engineagent.ts#1003), so reversing it silently drops the level. ThePOST /api/set-modelpath repeats the same model-then-effort order. A stale recorded level that the new model does not list is cleared by the composer on model switch (ticket 11 wire half). default_valuefromthinking_configis not a response field. The control's initial state is "Default" (thinkingPicker.none) until the user picks; for variant models an unpicked boot selects the engine's default variant (default_value: 'true'→ thinking on).- Engine-session entries appear in variant wire form (
m:...:v:thinking/:v:none-thinking) because that is what the engine advertises for switchable models; both carry the samethinkingLevels. - A pick while a turn is running takes effect on the next turn (same semantics as a model switch mid-run).
- A local pick owns its field for
PICK_DEFER_WINDOW_MS(4 s):applyConfigOptionUpdatedoes not overwrite that field with the engine's wire-formcurrentValueinside the window, so an optimistic pick is not clobbered a few ms later. Model and thinking are stamped independently (modelPickedAt/thinkingPickedAt), so a thinking-only pick does not block a later cross-client model mirror. The window is defence-in-depth — the per-cid snapshotrevisionis the primary guard against wire reordering. - An operator's providers-config entry with the same id as a builtin wins wholesale (existing merge rule); such an entry shows levels only if the operator wrote them.
- Ticket 49 batches 1–2 added a fourth display and a second editing entry inside the picker. The panel-bottom detail area renders the ACTIVE model's
thinkingLevelsas read-only badges. When a provider cascade is open, the fly-out renders as the reference picker's two-column popover — model rows on the left, a follow-focus settings column on the right — and that column's level control is editable and shape-adaptive: exactly["off","on"]renders one toggle switch; any other level list renders a radio group whose FIRST entry is always "Default" (submitted as the empty string). Form only — the wire semantics stay the localthinkingLevels+""contract, never the reference's effortOptions/variant derivation (A9 was scoped out). A pick commits immediately without closing the menu, through the same{thinking}payload the composer-level control sends; a focused-but-not-active model renders the control disabled with a "preview" marker. A recorded level the target model does not support highlights nothing — the row-badge anti-stale rule (B11) applied to the control.
The operator-facing view — which models show what control, and why MiniMax-M3 only has on/off — is the thinking section of webui.zh-CN.md.
Switching to another session re-points the active workspace to the
session's stored workspace. The file tree (slice 01) re-roots under
the new directory; the old project's webui:files-tree:<workspaceDir>
key is left intact in sessionStorage so the user returns to the
same expanded set if they switch back.
POST /api/sessions/switch (routes/sessions.js#handleSwitchSession)
resolves the workspace in this order:
- Target session's stored
workspace— that is the workspace the user was in when they last had this session open, modulo any pollution a previous code path introduced. DEFAULT_WORKSPACE(envMCODE_WORKSPACE> mcode TUIcwd.json>homedir()) when the stored value is empty. Empty is the shape legacy sessions or②-polluted records carry.- Refuse with
400if the resolved path failsassertWorkspacePathcontainment (e.g. the user tightenedMCODE_WEBUI_WORKSPACE_ROOTSsince the session was last opened).
The chosen path runs through the same assertWorkspacePath gate that
the workspace picker (/api/workspace),
browseWorkspace (/api/workspace/browse), the new-session POST, and
/api/fs/* all funnel through — refusing to switch into an
out-of-bounds path is the same boundary the picker refuses to land on.
The switch NEVER overwrites the target session's stored workspace
with the previous cs.workspace.dir. That was the pre-fix behaviour
(ticket 39's ② pollution path): every first-touch of an mvs_
session from project A inherited A's path, so the per-project
session grouping ended up duplicating A's directory for every
session the user opened from A. Newly-created overlays now start with
workspace: ""; the target-first read picks DEFAULT_WORKSPACE for
them.
The switch does NOT update cs.lastUsedWorkspace. The session bar's
"recent" sort (slice 07) is written only by handleSend; switching is
browsing, not authoring, and the previous "click any session and the
session jumps to the top" behaviour was the report that pinned that
contract.
Switching mid-run does NOT abort the in-flight turn. The run's
ownership and stream buffer are keyed by (cid, mcodeSessionId) in
lib/state-bus.js#runChatByCid, not by workspace. The engine child
process holds its own cwd from when the turn started; the new
cs.workspace.dir is purely the next-viewing surface. Finalize-time
behaviour (routes/chat.js finalize drain) is unchanged: still-viewing
writes into cs.chat; switched-away writes via
appendChatToSession(owningSid, lines).
A tab can run two conversations at the same time. The claim POST /api/send
takes is keyed by (cid, conversation), not by cid alone.
cid is the browser-TAB identity — one localStorage['webui_cid'], generated
once and deliberately stable across a session switch so one tab keeps one
client state object, one SSE channel, one coalesce/revision bookkeeping and
one mcode acp transport connection. Those stay tab-scoped on purpose. The
turn claim was not one of them: keyed by cid alone, a long turn in one
conversation refused every send into every other conversation of the same tab
with 409 cid-busy until it finished.
| Situation | Result |
|---|---|
| Send into session A while session B's turn runs (same tab) | 200, runs in parallel |
| Second send into the SAME session while its turn runs | 409 cid-busy — the duplicate-execution guard (#126 D-2) |
| Another tab or client already running that engine session | 409 session-busy |
More live turns server-wide than MCODE_MAX_CONCURRENT |
409 at-capacity |
The conversation key is the webui record id, with null as a first-class key:
it is the tab's unsaved draft, which is itself a conversation, and a tab has at
most one. handleSend claims before that record exists (a brand-new session has
no id yet) and creates it a few statements later with no await in between, so
the claim is re-pointed onto the new id by moveRunSession — otherwise the
second send into that conversation would find a free key and start a duplicate
turn.
Two consequences of the conversation being a first-class key are worth stating because they are load-bearing elsewhere:
-
A first turn's record id changes under the run. The draft is promoted to the engine identity mid-turn (
bindDraftToMcodeSid) andcs.sessionIdfollows, so the id a run was claimed under stops matching the view. Every lookup that answers "is this session the one that is streaming?" therefore falls back to the engine session id, which does not change. This is why a duplicate send into a first-turn conversation is answeredsession-busyrather thancid-busyonce the backfill has landed — both refuse. -
The claim moves with the id, and remembers where it was. The promotion is the one instant the conversation's identity changes, so
mcode-acp.jsre-keys the claim there (moveRunSession) on both transports. The registry keeps the retired key as an alias on the entry, which is what lets the route'sfinally { endRun(cid, runSessionId) }— still holding the key it claimed under — find and release the re-keyed claim.Without the re-key the guard has a hole, and the hole is about acknowledgement rather than about locking.
beginRuncannot see a turn whose key the view no longer presents, and its remaining guard (runsBySid) is populated by a separate mid-turn backfill. In a window where neither matches, the server answers200and hands a second concurrent turn to an engine session that is already executing — while the new turn's›echo lands in a livecs.chatthat the run-mirror's finalize then writes over from a snapshot taken before it. The result is the one failure this whole area exists to prevent: the engine ran the message and the webui holds no record of it, so the user gets neither the bubble nor the history entry and the text is gone. (Observed in the 16:00 UAT round, 2026-10-03, exception #1.) A guard that cannot see a turn must not ack it. -
MAX_CONCURRENTcounts turns, not busy clients. One tab running two conversations spends two of the slots, because that is two engine subprocesses; that is the resource the ceiling exists to bound.
A message sent into a conversation that is already running a turn is
refused, not queued. POST /api/send answers 409 with
reason: "cid-busy" or "session-busy", the turn is never handed to the
engine, the › line is never written, and nothing reaches the persisted
record. The refused text comes back to the composer.
The 409's error field is written for the person reading it — it names the
decision and the next action — because the composer renders it verbatim.
reason is the stable machine-readable key, and it is what the client
branches on rather than on the wording.
There is no queue, and the three send outcomes in the composer are kept distinct because they ask for opposite behaviour:
| State | What the server did | What the banner says | What the user should do |
|---|---|---|---|
| accepted | 200; the turn runs |
— | nothing |
| refused, conversation busy | 409 cid-busy / session-busy; the engine has nothing |
not delivered, text is back, wait for the turn | send again when the turn ends |
| unconfirmed | no answer, and the probe against the server could not establish whether the turn started | status unknown, or "the engine is running it, do not resend" | read the history first |
The third state is the one that must never lie about a side effect. It used
to treat "a turn is running" as proof that this send was accepted — the
reasoning being that a busy conversation answers 409 immediately, so a turn
seen after a deadline expiry is this one. That is false for the case that
actually produced the field report: the send was made into a running
conversation, so the running turn the probe sees is the previous one. The
banner then told the user "the engine is running your message, do not send it
again" about a message the engine never received. stateAcceptsSend now
requires the prompt's own echo line in the transcript, and consults the
running flag only when the snapshot carries no transcript at all — the one
place it cannot be contradicted, and where ignoring it is what made
sleep 35 execute twice under webui-parity 81 D-2.
What stays tab-scoped, and why it is safe under two live turns:
| Concern | Key | Why it is still correct |
|---|---|---|
| Client state, SSE channel, snapshot revision, push coalescing | cid |
One projection per tab is the contract; the snapshot is scoped per conversation by snapshotViewFields |
Stream line buffer (runChatByCid) |
(cid, engineSessionId) |
Already per conversation. createRunChat replaces only the calling run's own key — the previous whole-tab replace would have dropped a sibling's live lines |
| Run indicator / "thinking" state | the viewed conversation | viewOwnsLiveRun resolves the viewed session's run, so a sibling's turn neither claims nor clears this view's indicator |
Engine child process, /api/stop, session RPCs |
(cid, engineSessionId) |
One subprocess per turn. /api/stop and session/cancel / session/set_config_option target the VIEWED session's child; a tab-wide lookup would have signalled the wrong turn |
- Side effect: the file-tree panel re-roots under the new
workspace's root; expansion / filter / showHidden for the old
workspace stay preserved in
sessionStorage, the new workspace starts at its own stored expansion (or empty if never opened). - Failure mode: an out-of-bounds stored workspace returns
400with the gate's actionable message (工作区越界: <path> 不在任何允许根内。允许根: …).cs.workspace.diris NOT re-pointed; the previous workspace stays active. - Pre-existing polluted sessions (recorded before this fix carried the previous project's path): the new rule reads the stored value verbatim. The user sees the polluted path's file tree until they open the workspace picker and re-pick the intended directory; that one re-pick rewrites the stored value to the canonical realpath.
routes/sessions.js#handleSwitchSessionadds_resolveSwitchWorkspace(target, currentWs)and writescs.workspace = { dir: switchWs.dir, branch: null, tree: null }after the target is resolved and beforeresetContext.routes/sessions.js#handleSwitchSessionno longer passesworkspace: wstoensureOverlayForMcodeSid— new overlays start withworkspace: ""; the read picksDEFAULT_WORKSPACEfor first-touchmvs_switches.- The response payload gains
session.workspaceandsession.workspaceFallback; the SSE state push (handled bypushStateFor(cid)at the end) carries the newcs.workspace.dirverbatim, soFilesPanel'suseSessionContext()subscription re-renders without any client-side wiring change. routes/sessions.js#_eventsAppend("session.switch", …)recordsworkspaceandworkspaceFallbackso post-mortems can answer "why did the file tree jump".
The sidebar was aligned with the reference client across three areas:
collapse behaviour, nav rows, and the session list. This section is the
contract; docs/webui.zh-CN.md's sidebar section is the user-facing twin.
| Constant | Value | Where |
|---|---|---|
SIDEBAR_RAIL |
64 (px) |
components/shell.tsx — raised from 52 to match the reference rail |
SIDEBAR_MIN / SIDEBAR_MAX / SIDEBAR_DEFAULT |
240 / 400 / 240 (px) |
drag-resize clamp, unchanged |
SIDEBAR_AUTO_COLLAPSE_PX |
980 (px) |
innerWidth < 980 collapses; never auto-expands |
| Width transition | 180ms cubic-bezier(.2,.7,.2,1) |
on the outer wrapper, matching the reference's .webui-rail |
| Collapsed background | bg-transparent |
expanded keeps bg-bg_default_scrim |
| Collapsed toolbar compensation | pl-[142px] |
wraps the toolbar slot in AppShell, the reference's own constant |
The collapse toggle is a size-8 button on a pointer-events-none overlay
outside the sidebar (data-testid="sidebar-collapse-toggle", expanded at
left: 8px, collapsed at SIDEBAR_RAIL + 8), so it stays reachable in both
states instead of living inside the clipping card. Its aria-label flips
with state — sidebar.expand / sidebar.collapse, "Expand navigation bar" /
"Collapse navigation bar" (展开导航栏 / 收起导航栏) — and aria-expanded
carries the machine-readable half.
The collapsed form is an icon rail, not the reference's emptied strip:
rows and the avatar menu stay reachable. This is a deliberate keep — the
reference's own collapsed shell hides everything and its rail-mode
UserMenu branch never executes there.
Persistence: the collapsed flag round-trips through the existing
webui:ui-state payload (sidebarCollapsed, see lib/persist.ts
readShellCollapsedFromPersistedState / writePersistedShellCollapsed).
No new storage key was added, so the persistence-keys section is
unchanged.
Activation is a pure function, lib/sidebar-nav.ts#isSidebarNavActive, so
the regression suite drives the decision rather than rendered classes:
sidebar.searchlights when the tree column's active tab is thesearchsurface;sidebar.pluginslights for thepluginssurface (the legacypanel === "plugins"mirror agrees withopenSurfaceTab);topbar.newSessionlights only with no session selected and no sibling surface active — the reference's home-mode rule for its 新建任务 row.
The page computes the tree-surface signal (activeNavSurface, passed
through AppShell) because the tab strip state lives in page.tsx; the
session-id half comes from the store inside the shell. The active token is
bg_interaction_tertiary_hover held permanently — the reference's own
nav-active treatment, distinct from the session-row rule below.
定时 / 网站 / 远程 are not rendered: the repo's standing decision is that a control without a contract behind it is omitted, not shipped inert. They remain a product decision (see HANDOVER's 18094 fusion list).
- Selected rows (
sidebar-session-row,sidebar-subagent-row) paint withbg-bg_interaction_tertiary_selected; hover usesbg_interaction_tertiary_hover. Before ticket 47 both states wrote the hover token, which made the open session indistinguishable from any hovered row. Honest boundary: upstream defines the two tokens as the same value in the light theme (both resolve to--opacity_black_1_4, seetokens.css), so the visual distinction holds only in the dark theme (hoveropacity_white_0_4vs selectedopacity_white_0_8); the light-theme equality is the upstream token set's current state, not a regression introduced here. - Session and subagent rows are
<a href={sessionHref(id)}>deep links over the?session=query grammar (lib/url-restore.ts#sessionHref), not the reference's#session=fragment: the restore pipeline (cold load, popstate, replaceState sync) parses the query string, andwriteSessionToUrlpreserves fragments, so a fragment href would linger beside the query parameter. Plain left clicks are intercepted into the sameswitchSessioncall as before; modified clicks (middle / cmd / ctrl / shift) fall through to the browser, and the URL they open is one the cold-load path already restores. - Disclosures (project → directories → subagents) mount through
session-tree.tsx#Expandable, the reference's.webui-expandable-motion(grid-template-rows0fr→1fr, 180ms, plus a 140ms opacity crossfade);globals.csscarries the rule under the same class name with aprefers-reduced-motionbranch that drops the transition but keeps the open/closed state. Children stay mounted while collapsed (wrapped ininert+aria-hidden). - The reveal mechanism stays per-directory, 6 at a time
(
SESSION_VISIBLE_LIMIT); no global Load-more was introduced. GET /api/session-treeandGET /api/sessionsare read-only in this ticket — no request parameter or response field changed.
The reference has a right-click context menu (pin / archive / fork / delete)
and a localStorage pin+archive overlay; both need server contracts that do
not exist yet (no pin, archive or copy endpoints) and are deferred as a
batch. The Agent Team badge, the workspaceDir secondary line and the
「最近任务」 section exist in reference files but are not rendered by the
reference's own shell, so they are not implemented here either. A later
agent must not mistake any of these for "implemented but broken".
The conversation toolbar's right end of the title row carries a version badge: the branch name, the abbreviated commit id, and how long ago that commit landed. It answers "which checkout am I looking at" without opening a terminal, which is the question a user has when a build behaves unexpectedly and there is more than one checkout in play.
| State | What renders |
|---|---|
| Repository with at least one commit | branch + headSha + relative commit time |
| Directory that is not a repository | nothing — the whole element is absent from the DOM |
Repository with an unborn HEAD (git init, nothing committed) |
nothing; there is no commit to name |
| Detached HEAD | the sha alone, with no placeholder word standing in for a branch |
| Request failed or has not answered yet | nothing |
The absent cases are the contract, not an afterthought: an empty pill would be
a control that looks live and carries no information, so
resolveVersionBadge returns null and the component renders an empty
string. webapp/test/toolbar-version-badge.test.ts pins the rendered output
in both directions.
Placement. The badge sits at the far end of the title row (ml-auto),
opposite the session title it qualifies, and before that row's pr-20 reserve
— so the fixed right-4 launcher cluster can never overlap it. It follows
the running-turn indicator when one is showing. The relative time is the only
part with a narrow-width rule (hidden lg:inline): the branch name and the
sha are what identify a build, so they stay and the time gives way. Long
branch names truncate rather than pushing the bar wider.
The click copies the short sha. It is a real <button> with an
aria-label and a transient 「已复制」 confirmation, not decorative text. A
denied clipboard shows no confirmation rather than a confirmation the user
acts on.
Data and cost. One GET /api/git/status on workspace change — the same
endpoint the right-panel Git panel reads, not a second source of truth. It
does not poll: a git status on a large tree is a real index refresh, a
version identity changes when the user commits or checks out a branch rather
than on a schedule, and the Git panel already sets the precedent of fetching
on workspace change plus an explicit Refresh. The relative-time half needs no
refetch at all — it ticks off the toolbar's existing 1s ticker, which the
elapsed-timer already pays for. While the Git panel is open, two requests for
the endpoint are in flight; that is accepted rather than hoisting panel state
into a provider above the shell for a panel the badge does not render.
The model picker's settings detail renders at two levels (ticket 49 batch 2). The panel-bottom area always describes the ACTIVE model; when a provider cascade is open, the fly-out renders as the reference picker's two-column popover — the provider's model rows on the left, a follow-focus settings column on the right. Hovering or keyboard-focusing a model row switches that column to the model without picking it; a cascade that just opened (nothing focused yet) falls back to the active model, mirroring the reference. Both areas read the same draft mirror (below), so a window pick made in either place highlights in both.
A context-window radio group mounts only when the target model's /api/models entry carries at least two contextWindowOptions; a model without the field — or with a single option, which would be a no-op choice — renders no control. Today that is exactly MiniMax-M3 and MiniMax-M3.1-Flash-Preview ([512000, 1000000]); every other catalogue entry stays field-free. Each option label is a compact token count (512K, 1M), and an option the engine hints as higher_usage (contextWindowOptionHints) carries a "higher usage" tag. The two-column fly-out clamps itself to the viewport (max-height: 100vh − 16px) and each column scrolls its own overflow, so on short viewports the settings column's level control is never covered by the rows column.
When the focused row is not the active model, the settings column renders a preview: the options show (so the user can see what the model offers before picking it), but the controls are disabled and carry a "preview" marker — the recorded settings belong to the active model, and committing them for an unselected model has no contract meaning under the unchanged /api/set-model payload. Two empty states cover the rest: nothing describable (no focused row, no active model) shows "Select a model to see its settings"; a target model advertising neither context-window options nor thinking levels shows its MODEL NAME first, then "This model has no adjustable settings." — the container is an aria-live="polite" region, and announcing the bare sentence would leave a screen-reader user asking which model it is about. Where each control edits: a context-window radio commits from BOTH the settings column and the panel-bottom area — immediately, without closing the menu, so consecutive adjustments are possible; a LEVEL is pickable only in the settings column (or the composer-level control) — the panel-bottom area renders levels as read-only badges.
While the picker is open, the frontend also keeps a draft mirror (useState map keyed by model id): every pick lands in the mirror before the wire round-trip, so the highlight moves the instant the user clicks and does not blink back to the stale prop while the server confirms or /api/models re-fetches (a catalogue refresh mid-interaction never interrupts the flow). Closing the picker drops the mirror; reopening starts from the persisted state the server reported.
Close semantics (as measured): a pointerdown outside closes the menu; clicking a model row picks it and closes the menu; picking a setting commits and keeps the menu open (levels pick from the settings column; window radios pick from the settings column and the panel-bottom area alike); on the closed trigger, ↓/→ opens the menu and focuses the first row (↑ the last), and the panel's own engine takes over from there — → on a provider row opens the cascade and focuses its first model, ↑/↓ cycle inside the cascade, Home/End jump to its ends, so the whole chain is keyboard-only reachable; ← inside an open cascade closes only the cascade and restores focus to the provider row, leaving the menu open; Escape closes the whole menu (the cascade goes with it, and focus returns to the trigger).
Where the metadata comes from: the same engine-materialised builtin tree as the thinking projection (provider.minimax.models in <engine data dir>/config.yaml, keys contextWindowOptions / contextWindowOptionHints / limit.context). GET /api/models reads it on every request (readEngineBuiltinContextWindows, server/lib/engine-catalogue.js) and annotates both the builtin shell entries and the engine-session wire-form minimax_api entries; the highlighted value resolves to the recorded pick first and the model's contextLimit (the engine's current effective window) second, and is reported as currentContextWindow.
Contract note (ticket 49, both batches): neither batch changed any field of the /api/set-model request body — picks still travel as {model, contextWindow} / {contextWindow} / {thinking} exactly as U6 and the composer-level effort control shipped them (batch 2's in-picker level picks reuse the composer control's {thinking} payload verbatim), and the local thinking contract (thinkingLevels catalogue + "" = engine default) is untouched; the reference's effortOptions/variant derivation (A9) was scoped out and would be its own ticket with the server side. The provider grouping (models bucketed per provider with sticky headers) and the three thinking displays (row badge, chip suffix, composer control) are the two explicit user exceptions from PROMPT-ui-fidelity.md — they take precedence over anything the reference layout does and must survive any future picker rework.
Testability of those two exceptions: the derivations behind them are exported pure functions in packages/webui/webapp/lib/model-groups.ts — groupModelsByProvider, providerIdOfModel, isGroupDisabled, thinkingLevelsForModel, thinkingLevelKey, chipLevelSuffix, modalityBadgeKey, providerLabel — and components/composer.tsx imports that module rather than re-deriving the rules inline. webapp/test/composer-models.test.ts therefore drives the product functions; it previously re-implemented the grouping loop in the test file, which made red line ⑤ unfalsifiable (a broken grouping stayed green). The extraction is behaviour-neutral: the code moved with its inputs named, and nothing about what a user sees changed. webapp/test/composer-thinking-tripwire.test.ts keeps pinning the call site, so a half-done extraction — an exported function the selector no longer calls — fails.
Honest boundary — a pick is recorded, not yet engine-applied. POST /api/set-model accepts contextWindow (tokens; null clears), validates it, records it in cs.model.contextWindow, and echoes it in the response. The engine's ACP surface has no channel for it: session/set_config_option accepts exactly three config ids, and the model value's wire encoding (m:<provider>:<model>:u|v:<variant>, packages/tui control-state.ts#modelConfigValue) has no context segment — verified against the shipped engine bundle (0.5.5) as well as this repo's source, whose runtime models.select does accept a contextLimit but is reachable only from the TUI/runtime clients. The recorded pick is therefore a webui-side preference the picker reflects immediately; the model switch that always accompanies it does reach the engine through the existing set_config_option{configId:"model"} push. Wiring the value into an engine-side apply is the engine ticket's work, and the route's shape (validate → record → echo) is the seam it plugs into. The same follow-the-model rule as thinking applies: switching to a model that does not list the recorded window clears it (contextWindow: null) in the same request.
Every shipped file tree, panel and column evidence is grep-able. The list
below cites the component file and one data-testid per surface.
| Surface | Component | Anchor data-testid |
|---|---|---|
| Sidebar (rail) | components/shell.tsx |
sidebar-scroll-viewport |
| Sidebar collapse toggle (overlay outside the rail, ticket 47) | components/shell.tsx |
sidebar-collapse-toggle |
| Sidebar session tree | components/session-tree.tsx |
sidebar-session-row |
| Session-tree section header (plain text, ticket 47) | components/session-tree.tsx#SectionHeader |
sidebar-section-header |
Session-tree error state (role="alert", ticket 47) |
components/session-tree.tsx |
sidebar-tree-error |
| Sidebar user menu (settings / upgrade / check-in / usage / feedback & help / sign-out + trailing user card; full row set since ticket 55c) | components/shell.tsx#SidebarFooter |
sidebar-user-menu |
| Project context menu (ticket 55c) | components/session-tree.tsx#ProjectNode |
project-context-menu |
| Home quick-capability capsules (ticket 55c) | components/chat.tsx#HomeState |
home-quick-capabilities |
| Sidebar inbox (alerts flyout) | components/inbox.tsx |
inbox-flyout |
| Toolbar (top bar with model selector) | components/toolbar.tsx |
toolbar-session-status |
| Toolbar version badge (branch + short sha + commit time, webui-parity 89) | components/version-badge.tsx |
toolbar-version-badge |
| Composer + drop overlay | components/composer.tsx |
composer-drop-overlay, composer-send-button |
| Chat (virtual list ≥ 200 messages) | components/chat.tsx + chat-virtual-list.tsx |
chat-virtual-top-spacer |
| Turn process bar (composite summary + output rate since ticket 46 PR3) | components/activity-group.tsx#TurnProcessDisclosure |
turn-process-disclosure |
| Turn-bar expand chevron (webui-parity 61, drives the turn's activity groups) | components/activity-group.tsx#TurnProcessDisclosure |
turn-process-trigger, turn-process-chevron |
| Streaming-label phrase rotation (webui-parity 61) | components/loading-states.tsx#ActivityPulse + lib/thinking-phrases.ts |
activity-indicator-label |
Activity group (collapsible tool turns; in activity-group.tsx since ticket 46) |
components/activity-group.tsx |
activity-group-header |
| Thinking block (thought-process disclosure row, ticket 46 PR2) | components/activity-group.tsx |
thinking-block |
| Tool card (one tool call, ticket 46 PR3) | components/activity-group.tsx#ToolCard |
tool-card |
| File preview (right preview column body) | components/file-preview.tsx + file-preview-pane.tsx |
file-preview |
| File tree column (column 4) | components/workspace-tree-column.tsx + panels.tsx#FilesPanel |
files-tree-root |
| File tree search (server-driven, slice 19a; wired in 19b) | components/panels.tsx |
files-tree-filter |
| Sidebar tree-column "搜索" surface (slice 19b) | components/workspace-tree-column.tsx#SearchSurface |
tree-surface-search-input |
| Code preview (slice 22 IDE-grade: gutter + lazy hljs + byte-faithful copy) | components/code-view.tsx |
code-view (rendered inside file-preview) |
| Three-state appearance picker (slice 18; lives in the settings General section since ticket 37) | components/appearance-card-picker.tsx |
appearance-card-picker |
| Git panel (slice 03) | components/panels.tsx#GitPanel |
git-panel |
Browser panel (slice 04, sandboxed iframe over /api/fs/raw) |
components/browser-panel.tsx |
browser-panel |
| Workspace picker (modal) | components/workspace-picker.tsx |
workspace-picker |
| Provider management | components/provider-management.tsx |
providers-panel |
| Add-model dialog + fetched-models dialog (ticket 54; lifted to its own file in acceptance round 2 so the suite can render-test it; ticket 56 visual parity) | components/add-model-dialog.tsx#AddModelDialog / #FetchedModelsDialog (controlled surfaces #AddModelDialogForm / #FetchedModelsDialogBody, pure helpers #collectDialogErrors / #defaultChecked) |
provider-dialog (fields provider-dialog-provider-select / -api-key / -api-key-reveal / -model-add / -autofetch / -models-empty / -cancel / -save / -footer / -errors; per-entry provider-dialog-entry-{n} with -name / -context / -max-output / -thinking / -attachment-{mod} / -test / -test-result / -reset / -remove) / fetched-models-dialog (fetched-models-title / -item-{id} / -select-all / -cancel / -add) |
| Context meter / panel | components/context-meter.tsx |
context-meter |
| Settings modal | components/panels.tsx#SettingsModal |
settings-modal |
| Segmented tabs of the Usage & models section (ticket 53) | components/panels.tsx#UsageModelsSection |
usage-models-segment (tabs usage-models-tab-token-plan / usage-models-tab-custom-models) |
| Plan / usage / credits / invoice cards of the Usage & models section (ticket 37, reworked 53) | components/panels.tsx#PlanCard / #UsageCard / #CreditsCard / #InvoiceCard |
settings-plan-card / settings-usage-card (bars usage-bar-fiveHour / -weekly / -video) / settings-credits-card / settings-invoice-card (invoice-apply-link) |
| Error boundaries (global + per-route) | app/error.tsx + app/global-error.tsx |
global-error-page |
On the right side of the sidebar, the shell renders a flex row that can
hold up to three visible columns: conversation | preview | tree. The
sidebar is owned by AppShell, lives outside the row, and is allocated
zero width inside the wrapper (workspace-tabs-state.ts:738-744).
Every width number below comes from COLUMN_SPECS in
packages/webui/webapp/lib/workspace-tabs-state.ts:412-456. Change a
default here, and the corresponding DEFAULT_COLUMN_LAYOUT follows;
change a minWidth / maxWidth, and clampWidth
(workspace-tabs-state.ts:511-515) plus computeColumnLayout
(workspace-tabs-state.ts:724-854) pick it up on the next render.
| Column | Role | Default | Min | Max | Flow |
|---|---|---|---|---|---|
sidebar |
AppShell chrome — outside the column row | 240 | 220 | 400 | fixed |
conversation |
Elastic — absorbs leftover, never caps the growth path | 720 | 280 | 2400 | fluid |
preview |
On demand — viewing surface (file:<path>, browser) |
400 | 320 | 720 | fixed |
tree |
On demand — navigation surface (files, git, tasks, search, plugins) |
340 | 320 | 600 | fixed |
The only maxWidth that does not cap what the user actually sees on
screen is conversation.maxWidth = 2400. It bounds the value the user
can write into the layout by dragging the divider
(workspace-tabs-state.ts:511-528); the layout algorithm explicitly
ignores it and lets conversation absorb every leftover pixel after
the fixed columns have grown to their caps
(workspace-tabs-state.ts:805-823, commentary at :438-443 and
:702-720). COLUMN_SPECS.conversation.maxWidth is therefore not
a viewport ceiling — past 1920 the column just keeps growing, and the
readable measure on the chat content (960 px, centred —
components/chat.tsx:38-52, applied at :255 and composer.tsx:532)
takes over the visual constraint.
computeColumnLayout (workspace-tabs-state.ts:724-854) is a pure
function that returns a ColumnLayoutSummary. The flow is:
- Each column starts at its stored width (clamped to its
[min, max]band).conversation's stored width is the user's drag target;preview/treestart at 0 when collapsed (slice 21 on-demand model —workspace-tabs-state.ts:494-504). - Overflow → fixed columns shrink toward their minimums in the
order
preview → tree. If still over,conversationshrinks toward its minimum (280). Last resort:conversationshrinks below its minimum and the renderer hides it. - Leftover → fixed columns grow toward their maximums in the
order
tree → preview(the fold order reversed). Whatever is left after that goes toconversation, unconditionally — slice 25 removed the previous growth-path ceiling, and the source comment pins the decision (workspace-tabs-state.ts:805-823).
Two invariants follow directly:
- Row sums to exactly the container width, every render (modulo zero-width hidden segments). No dead gutter is reachable.
- When
conversationis the only visible column, its rendered width stays ≥ 280 px whenever the container can fit 280 px — the fold ladder in step 2 shrinkspreviewandtreefirst, soconversationonly shrinks past 280 when nothing else gives. At very narrow viewports the renderer hides it entirely (workspace-tabs-state.ts:762-781).
When both preview and tree are closed (the shipped default on first
paint — workspace-tabs-state.ts:494-504), conversation absorbs
every pixel the row has after sidebar has taken its default 240 px.
The conversation width is therefore viewport − sidebar, a computed
remainder, not a configured cap:
| Viewport | Sidebar | Conversation (idle) | Notes |
|---|---|---|---|
| 1280 px | 240 | 1040 | 1280 − 240 |
| 1920 px | 240 | 1680 | 1920 − 240 |
| 2560 px | 240 | 2320 | 2560 − 240 |
The chat content's 960 px centred measure (components/chat.tsx:38-52,
applied at :255 and composer.tsx:532) is a hard CSS cap on the
content, not a threshold on the column. As soon as the column reaches
960 px the cap starts biting: the content stays at 960 and the slack
above it splits evenly between left and right inside the column. The
column itself keeps absorbing leftover up to the algorithm's only
limit, which is the container width.
When at least one fixed column is open, leftover after the user's
stored widths flows into the fixed columns first, bounded by their
maxWidths (workspace-tabs-state.ts:793-815). At a 1920 px viewport
with both fixed columns open at their defaults, the container is
1920 − sidebar 240 = 1680 — the same 1680 figure the idle table
and the "only preview" example use (workspace-tabs-state.test.ts:1028).
The stored widths (preview 400, tree 340, conversation 720) sum to
1460, so:
- container leftover =
1680 − 1460 = 220 treegrows 340 → 560 (eats 220 px, falling short of its 600 max because that is all the leftover there is)previewstays at 400 (no leftover left)conversationends at its stored 720 — no leftover reaches it
If only preview is open at 1920, tree is skipped (collapsed). The
container is still 1680; the stored widths sum to 1120, so the
leftover is 560:
previewgrows 400 → 720 (its max, eats 320 px)conversationabsorbs the remaining 240 px on top of its stored 720 → ends at 960 (the locked value inworkspace-tabs-state.test.ts:1077).
Drag behaviour on conversation itself is bounded by [280, 2400] via
clampToConversation (workspace-tabs-state.ts:856-861); the
algorithm may then re-distribute any overflow into preview first
(workspace-tabs-state.ts:747-781).
Each column hosts its own independent activeId (previewActiveId,
treeActiveId) so opening a tree surface does not steal focus from the
preview column, and vice versa. The surface vocabulary
(SurfaceTabKind) is six values — files | git | tasks | search | plugins on the tree side, browser | file:<path> on the preview side
— and is the single source of truth in
lib/workspace-tabs-state.ts#SURFACE_TAB_KINDS. The sidebar's
"搜索" surface is real as of slice 19b — the
SearchSurface component in workspace-tree-column.tsx wires a
200 ms-debounced request to GET /api/fs/search (api.searchFs),
reuses the same searchFootSegments footer as the file-tree filter
(scanned / matches / skipped / truncated / budget), and on click
sends an expand-to-hit request through the shared fs-tree-reveal
channel so the file tree panel applies the same expand + highlight.
The plugin panel shipped with the plugin backend (ticket 68).
PluginsSurface (webapp/components/plugins-surface.tsx) is mounted in
both column hosts (panels.tsx:354, workspace-tree-column.tsx:645), so
the ten /api/plugins/* endpoints in
The plugins API below are called by
a screen a user can open today. The plugins area is real data — the
installed list, the local marketplace, GitHub import — and in the
installed view every card carries an enable switch and an uninstall
behind a confirmation dialog. The skills, apps, mcp and agents
areas have no management endpoint yet, so they render a pending card
naming the capability that is missing rather than a list of invented
entries.
Surface kinds go through openSurfaceTab("…"); the right-panel kinds
(PanelKind) are a separately-trimmed union: "workspace" | "files" | "git" | "plugins" | "browser". The previously-shipped search, alerts,
and progress kinds have been removed from the PanelKind union
(packages/webui/webapp/lib/persist.ts#PanelKind); alerts is reached via
the separate bell-icon InboxFlyout component, and progress had no live
entry point at all.
Dividers between columns are 8 px wide and support drag-resize (clamped
to [minWidth, maxWidth] per column) and double-click reset.
This section documents the contract behind that panel: the ten
endpoints, their parameters, and the states a caller must handle. What
the user actually gets in the plugins area is a market view and an
installed view, a keyword box on both, a category dropdown and a source
switch on the market side, a refresh button, and a two-step GitHub
import (preview the URL, then commit it). Only the plugins area loads
anything; the four pending areas never issue a request. Two data facts
decide how much of that is reachable: the local marketplace, the
installed list and GitHub import return real data, while the official
marketplace cannot resolve in a local edition.
| func_name | Endpoint | api.ts function |
Parameters |
|---|---|---|---|
plugins.list.installed |
GET /api/plugins/installed |
listInstalledPlugins |
keyword? limit? cursor? |
plugins.list.marketplace |
GET /api/plugins/marketplace |
listMarketplacePlugins |
source (required) + the above, plus category? skillLimit? skillCursor? |
plugins.list.enabled |
GET /api/plugins/enabled |
listEnabledPlugins |
— |
plugins.refresh.all |
POST /api/plugins/refresh |
refreshPlugins |
— |
plugins.enable.by_name |
POST /api/plugins/enable |
enablePlugin |
pluginName source? |
plugins.disable.by_name |
POST /api/plugins/disable |
disablePlugin |
pluginName source? |
plugins.install.by_name |
POST /api/plugins/install |
installPlugin |
pluginName source? |
plugins.uninstall.by_name |
POST /api/plugins/uninstall |
uninstallPlugin |
pluginName source? |
plugins.import.preview_url |
POST /api/plugins/import/preview |
previewGithubPlugin |
url |
plugins.import.from_url |
POST /api/plugins/import |
importGithubPlugin |
source (repositoryUrl commitSha subPath?) |
webapp/lib/api.ts exposes one typed function per endpoint:
States a caller must handle, as the contract defines them. The rendering
column is what the shipped panel does, with one exception:
plugins.refresh.all is never called — the panel's refresh button
re-pulls the list rather than posting to the endpoint.
| func_name | empty | loading | error | success |
|---|---|---|---|---|
plugins.list.installed |
{ok:true, plugins:[], hasMore:false} |
consumer's own | 200 {ok:false, code} |
one page, official + local merged |
plugins.list.marketplace |
{ok:true, plugins:[]} |
consumer's own | source=2 → report it; source=1 → designed notLocal state |
plugin rows + marketplaceSkills for the local source |
plugins.list.enabled |
{ok:true, plugins:[]} |
— | 200 {ok:false, code} |
{plugins:[{name, displayName?}]} |
plugins.refresh.all |
— | refresh-button spinner | runtime code passed through | {ok:true}, then re-pull the installed list |
plugins.enable.by_name |
— | row spinner | PLUGIN_NOT_FOUND / PLUGIN_AUTH_REQUIRED / PLUGIN_AUTH_SYNC_TIMEOUT |
{ok:true, sourceKind, installExists, enabled:true} |
plugins.disable.by_name |
— | row spinner | same three codes | {ok:true, sourceKind, installExists, enabled:false} |
plugins.install.by_name |
— | button spinner | PLUGIN_AUTH_REQUIRED; LOCAL_PLUGIN_INSTALL_UNSUPPORTED on a local package |
{ok:true, sourceKind, installExists:true, enabled:true} |
plugins.uninstall.by_name |
target absent → {ok:true, installExists:false} |
confirm dialog, then spinner | same three codes | {ok:true, sourceKind, installExists:false, enabled:false} |
plugins.import.preview_url |
— | dialog loading | invalid URL / PLUGIN_NO_SUPPORTED_CAPABILITY / unreachable |
{source, plugin:{summary,…}, diagnostics, packageSizeBytes, canImport} |
plugins.import.from_url |
— | button spinner | PLUGIN_ALREADY_EXISTS / PLUGIN_IMPORT_INVALID |
{plugin:{summary}}, already enabled |
The one endpoint that cannot serve a local edition, stated plainly.
The official marketplace needs a cloud account, and the local edition's
cloud base URL does not resolve, so source=1 answers
{ok:false, code:"NETWORK_ERROR"}. The panel treats that as a designed
state — the plugins.market.official.notLocal.* copy rather than a red
error — and short-circuits before issuing the request
(mayRequestMarketplace), because a request that fails only after a
30 s timeout would make a designed state look like an incident. Official
refusals on the install / enable / disable / uninstall actions are
silent for the same reason. Everything else is real
data: the installed list, the local marketplace (standalone skills plus
the local package projection), and the two GitHub import endpoints, which
fetch a public repository directly and never touch the registry. A
local package cannot be installed — the runtime answers
LOCAL_PLUGIN_INSTALL_UNSUPPORTED — so canInstall is true only for an
official row in the market view and a local card renders no install
button, rather than offering an action that always fails.
Wire conventions worth knowing before writing a new call.
source is numeric on the way in (1 official, 2 local); on the way
out the runtime's number is passed through and the route adds a
protocol-free sourceKind string ("official" / "local") to the
page, to every plugin row and to every mutation answer. A consumer must
branch on sourceKind, never on the numeric source, which is how the
webapp stays free of an @mavis/protocol dependency. The marketplace
listing requires source, because the runtime reads a missing one as
"official" and a silent default would send every request to an
unreachable registry. A
runtime failure answers 200 with ok:false and a code to branch on;
a rejected request answers 400 instead, and the webapp helper surfaces
any non-2xx as a thrown error carrying the server's message — which
means a code on a 400 cannot be read back, so a consumer must reset a
stale cursor on a filter change rather than parse the failure. A cursor
is bound to the filter it was issued for: reusing it after a keyword
change answers 400 PLUGIN_CURSOR_INVALID. Authentication is the shared
gate chain, and in
read-only mode every POST answers 403 — the mutations are unavailable
by policy, not by failure.
The file preview tab uses components/code-view.tsx (data-testid
file-preview). It layers three slice-22 affordances on top of the
plain <pre> view that shipped in slice 02:
- Line-number gutter, aligned to code lines and independent of
horizontal scroll — line numbers never move when the user scrolls
right on a long line.
splitHighlightedLines(webapp/lib/code-highlight.ts) walks the highlight.js HTML output and balances any<span>that crossed a line boundary, so each line is hover-stable and copy-faithful. - Per-language lazy syntax highlighting. The grammar for the open
file's language is the only grammar loaded —
loadHljsLanguageis a switch / if-ladder of literalimport("highlight.js/lib/languages/<name>.js")branches so webpack code-splits each grammar into its own chunk (the alternative — a Record-driven dynamic import — would have bundled all 191 grammars). Unknown or unloaded languages fall through to a plain monospace view (the contract is total: bad inputs must not blow up). The hard byte cap is 32 KiB with a 1500-line cap; larger files are truncated before the highlight step so a multi-megabyte file cannot freeze the tab, and the UI renders an honesttruncatednotice. The map is inLANGUAGE_TO_HLJS—htmlis an alias ofxml,jsoncsharesjson, andtoml/plaindeliberately have no entry (the caller treats them as plain monospace). - Byte-faithful copy. The copy path restores the trailing
newline (
endsWithNewlineis tracked across the highlight → split → copy chain so the clipboard text round-trips to the file bytes forcp file.js file.js.bak; copy in panel; paste back) and never leaks the gutter line numbers into the copied text.
The settings surface is a full-viewport modal — a grouped category tree on
the left (with a search box), a content column on the right (840px for the
General page, 760px for every other tab since ticket 48; the previous flat
704px cap had copied the usage page's own width). It is opened from the
avatar's user menu at the bottom of the sidebar. This section records the
structure and the write paths; the parity reference is the desktop's own
settings (refs/ui/03-settings-usage-models.jpg, 04-settings-general.jpg).
Navigation and capability honesty
Ten tabs in four groups. Every tab carries the reference's 18×18 stroke glyph; the state column says what a user actually gets, and a control that renders but cannot act is called placeholder — a designed outcome, not a missing feature. None of the ten is a bare placeholder any more: the last one, Worktree, reads the engine since PB-3.
| Group | Tab | State |
|---|---|---|
| Preferences | General (通用) | implemented |
| Preferences | Voice | implemented, placeholder controls — the microphone dropdown is disabled with a single 「本地版不适用」 option, and both dictation rows show 未设置 (no device enumeration, no dictation input in a browser) |
| Preferences | Shortcuts | implemented — 10 desktop rows, each stating what the browser can do with it: 3 rebindable and live, 1 live on macOS only, 6 blocked with the specific reason (see Shortcuts — what the browser can intercept) |
| Preferences | Personalization | implemented, and honest about what it is for — 自定义指令 and 关于你 persist to localStorage and each field states 「已保存于本浏览器,不会注入引擎会话」, because the engine has no channel that reads them (SB-8 / D-2; the storage is kept, the injection claim is not); both memory switches render off and disabled with the not-applicable marker, and 管理 opens the 记忆摘要 dialog in its permanent empty state |
| Management | Usage & models | implemented; since SB-1 the two engine sources are real (Token Plan / MiniMax API switch the engine's credential, the 「使用中」 badge reads the engine back, and the MiniMax API key can be saved and probed) — the third pill, Custom models, stays a VIEW onto the provider catalogue |
| Management | Connection | implemented |
| Management | Account | implemented as a read — the section reads GET /api/account on mount and renders the account name, the current plan name, the quota overview (plan-quota state plus the 5-hour and weekly remaining figures) and the account status; sign-out stays disabled (no engine method acts on it) |
| Coding | Code review | implemented — 自定义审查准则 persists to localStorage; 审查方式 is a disabled single-option dropdown showing 子会话 |
| Coding | Worktree | implemented as a CLEANUP page since PB-3 — see Worktree — what the page can and cannot do |
| Coding | Code review | implemented — 自定义审查准则 persists to localStorage and carries the same 「不会注入引擎会话」 note as the two Personalization texts (SB-8 / D-2); 审查方式 is a disabled single-option dropdown showing 子会话 |
| Coding | Worktree | not implemented — the tab is a one-line panel reading 「本地版暂不支持工作树管理」 |
| Archived | Archived tasks | the tab renders its empty state 「暂无已归档任务」; the list and its actions need an archived-session contract that does not exist |
Worktree — what the page can and cannot do. The Worktree tab is a
cleanup page, not a workspace manager. It lists the worktrees Git already
knows about for the current repository and removes the ones you select; the
desktop reference (design-ref/screenshots/ref-23.jpg) has no 「新建工作树」
button, and ManagedWorktreeServicePort declares no create either, so adding
one would be inventing a capability on both sides at once. Two endpoints:
| Endpoint | Engine call | Notes |
|---|---|---|
GET /api/worktrees?workspace=<repo> |
services.managedWorktrees.list |
workspace is optional; without it the request's own conversation workspace is used, and with neither source the endpoint is 400 rather than guessing the server's cwd |
POST /api/worktrees/remove |
services.managedWorktrees.removeBatch |
body {items: [{workspace, worktreeDir}], activeWorktreeDir?} |
Both are gated by PB-8's three-state presence gate rather than by a capability
key, because no key covers services.managedWorktrees: no host → 503
engine_host_unavailable, a host with no owner graph → 501
engine_services_unavailable, an owner graph without the service → 501
worktree_service_unavailable. None of them answers 200 with an empty list — a
page that says "nothing to clean up" for a runtime that failed to boot would
tell the user their project is clean when it was never read.
A folder that is not a Git repository is a report, not a failure: 200 with
ok: false and the engine's own code (not_git_repository,
workspace_unavailable, worktree_list_failed). The page names the code, so
the three mean three different operator actions instead of one 「读取失败」.
The three toolbar tabs (近 3 天 / 3-7 天前 / 7 天以上) filter on
lastModifiedMs and the boundaries are inclusive at the top of each band:
exactly 3 days old is still 近 3 天, exactly 7 days old is still 3-7 天前. They
filter rather than sort, because each tab is an age band in the desktop's mental
model, and sorting would make the older bands unreachable without scrolling.
A row whose timestamp the engine could not read (lastModifiedMs is genuinely
optional — the engine falls back from the directory mtime to the last reflog
entry) appears in every tab labelled 时间未知. Filing it at 0 would put a
worktree modified seconds ago under 7 天以上; hiding it would make a real
worktree invisible.
removeBatch verdicts pass through per item and are never collapsed: a batch
where every item was refused is ok: true with a full failedItems list, each
carrying its WorktreeRemovalReason (main_worktree / active_worktree /
not_found / locked_worktree / dirty_worktree / unknown). The page maps
each through one table and lists them under the toolbar. The main worktree, a
locked worktree and the worktree an active session runs in render their
checkbox disabled with the matching reason beside it, because a checkbox
that ticks and then fails on submit teaches the user the button lies. The
runtime-safety check the engine runs before any removal
(listRunningWorktreeDirs) is inside the service, so the UI cannot bypass it.
The workspace a browser names is containment-gated on both endpoints
(assertWorkspacePath, the same boundary as /api/fs/*); a batch is refused
as a whole when one of its repositories is out of root, because an out-of-root
repository is a forged request rather than a worktree that happened to fail.
worktreeDir is deliberately not gated separately: the engine only removes
a path that git worktree list reports as a linked worktree of that repository,
which is strictly stronger than a root check, and adding the weaker gate in
front would make legitimate out-of-root worktrees (a sibling checkout next to
the repo) unremovable.
There is no Browser tab in Settings. The browser surface is a workspace
column tab (workspaceTabs.tab.browser) that mounts BrowserPanel over the
workspace tabs, not a settings section; the settings.tab.browser dictionary
key has no call site. An earlier revision of this document listed a Browser
tab under Preferences.
Shortcuts — what the browser can intercept
A page cannot register a global shortcut, and it cannot intercept a
combination the browser has already claimed. The Shortcuts tab therefore
does not present the desktop's bindings as a dead reference list: every
row states its own verdict, and the same registry decides both the label
and the dispatch. webapp/lib/shortcuts.ts holds that verdict; app/page.tsx
matches keydowns against it and components/settings-extra-pages.tsx
renders it, so a row cannot be shown as live while nothing dispatches it
(or the reverse).
| Row | Combination | Verdict | Why |
|---|---|---|---|
| 显示或隐藏 Mini Chat | Alt+M |
blocked | the WebUI has no Mini Chat surface |
| 全局搜索 | Ctrl+K |
live | unclaimed in the browsers this client targets; opens the tree-column search surface |
| 搜索任务和会话 | Ctrl+G |
blocked | the browser's find-next |
| 新建任务 | Ctrl+N |
macOS only | a new window in Chromium and Firefox on Windows and Linux — the keydown never reaches the page there |
| 新建无项目任务 | Ctrl+Alt+O |
live | unclaimed |
| 打开项目文件夹 | Ctrl+O |
blocked | the browser's Open File dialog |
| 打开设置 | Ctrl+, |
live | unclaimed |
| 按住听写 / 切换听写 | — | blocked | no speech recognition behind the rows |
| 反转跟进行为 | Ctrl+Enter |
blocked | the key is free, but the action's semantics are undecided; binding it would promise behaviour that does not exist |
The three live rows are rebindable: the box takes the next combination
the user presses, persists it under webui-shortcut-bindings, and the
handler picks it up on the next keydown. A combination another dispatched
row already owns is refused and the conflicting action is named — two rows
sharing one combination would be an order-dependent bug in the handler.
Ctrl+N is deliberately not rebindable: moving it would not make it fire
on the platforms where the browser owns it, so the row prints the limit
instead of pretending a rebind fixes it. Blocked rows keep the desktop's
printed combination for reference, render disabled, and print the reason
from the table above. Ctrl+Shift+T-style interception is not attempted
and cannot be: those keys never arrive.
General (通用) — sections
The General page follows the reference's sectioned layout (ticket 48): an
<h3> title above each 16px-radius card, horizontal rows (title +
description left, control right, 56px min height) with hairline dividers
between adjacent rows:
| Section | State | Notes |
|---|---|---|
| Mode | disabled furniture | the two mode cards (built for coding / built for everyday work) render as the reference's, both disabled with the coding card preselected — no local mode switch |
| Application | enabled | appearance picker + language switch. The reference's five desktop switches (menu-bar icon, launch-at-login, desktop notifications, early access, accelerated indexing) render disabled, one per row — no capability behind them |
| Link destinations | disabled furniture | two rows (web links, local links) whose selects are disabled single-option dropdowns |
| Files | enabled | two switches persisted in localStorage, see the table below |
| Session management | enabled | one switch, persisted; gates the composer's context-window readout (see below) |
| Agent control | disabled furniture | the 「自动打开浏览器面板」 switch renders off and disabled (no capability behind it) |
| Preference settings | enabled | follow-up behaviour (disabled / queue / send now); since SB-4 the composer reads it, and a send into a running turn reaches the engine's queue or steers the running turn. Watermark and data opt-in render disabled |
| About | mixed | export logs is a live download of this server's own diagnostic trail (GET /api/logs/export, see below); check-for-update is still a disabled button — self-hosted update is git pull, and the desktop updater's semantics do not apply; the local URL and LAN URL are live read-only rows from /api/settings |
| dataDir footer | not implemented | the reference prints the app data directory at the bottom of the General page; /api/settings has no such field and the server routes are read-only this round, so no value exists to print |
Appearance and language behave as before: immediate effect on click; the
appearance choice persists through the webui:ui:v1 envelope; in system
mode the page follows the OS colour scheme live.
Browser-local switches (ticket 48)
Four bare-string localStorage keys, same names and format as the desktop
reference, so one browser profile carries the same preferences in both
clients:
| Key | Default | Affects behaviour? |
|---|---|---|
file_open_in_new_tab |
"true" here ("false" in the reference) |
Yes. On (default) keeps this client's standing one-tab-per-file behaviour; off replaces the active file tab with the newly opened file. The strip has no pinned-tab concept, so "active file tab" is the reuse target — a documented approximation of the reference's "reuse the unpinned tab" |
file_line_wrap |
"true" |
Yes. On wraps over-wide lines; off scrolls horizontally. Covers both code-file previews (ticket 48) and markdown codeblocks — chat messages, activity groups and markdown file previews (ticket 52); the language label never wraps. Applies to previews opened / messages mounted after the switch (an already-open one does not reflow); a wrapped file-preview line's gutter number aligns with its first visual row — a known trade-off |
webui-context-window-usage |
"false" |
Yes. On draws the context-window readout in the composer's toolbar, immediately left of the model chip; off renders nothing there. The readout's own form is unchanged — the ring, the percentage, the breakdown and the plan rows all come from the session snapshot as before. Flipping the switch takes effect without a reload |
webui-follow-up-behavior |
"queue" (or "off", "steer") |
No. Decides what a send does while a turn runs; "off" is webui's own third position (the reference has two) |
webui-desktop-notifications |
"false" |
Yes. On raises a browser notification when a turn finishes, when a tool needs a decision, or when a turn fails — and only while this page is in the background. See Desktop notifications (SB-9) below |
Desktop notifications (SB-9)
The 桌面通知 row on the General page is the one switch in the 应用 block that
is neither an OS integration nor a placeholder. The decision about whether
to notify is a pure function over four facts in
webapp/lib/desktop-notify.ts#shouldShowDesktopNotification — the stored
switch, the browser's live Notification.permission, whether the browser
supports the API at all, and document.visibilityState. The browser calls
live at the edges of that module and nowhere else.
| Question | Answer | Why |
|---|---|---|
| Which events notify? | turn settled, a needs_authorization request arrived, an error-level anomaly landed on the alerts stream |
Each already arrives on a stream this client subscribes to, so no new server channel was added. Subagent completion and plan review are deliberately not separate triggers: both arrive as needs_authorization, and a fourth trigger would be a fourth thing to get wrong |
| When is a notification suppressed? | whenever document.visibilityState === "visible", for all three kinds |
A notification that duplicates something already on screen is noise. Each of the three has an on-screen face: the finished conversation is in the tab in front of the user, the decision prompt is a blocking modal, and the failure is already in the alerts badge and the action-error banner. The reference has no such rule because it has no page to be looking at |
| What happens when a turn fails? | one notification, not two | The error and the settle arrive on two different SSE connections (alerts vs. events) and nothing orders two HTTP responses, so either order is possible. The error is remembered against its session and a completion on that session inside a 5 s window is swallowed; outside the window the stale entry is ignored, so it cannot suppress an unrelated turn later. The entry is consumed on use |
| What about a tab opened over a busy server? | the alerts opening snapshot is history, not news | GET /api/alerts opens with a ring buffer of up to 200 pre-load alerts. lib/alerts.ts#historySealed reports whether that frame has arrived, and until it has nothing in the buffer is announced. Told by the frame's own kind, never guessed from whether the list happens to be empty — an empty history and an unwatched first failure are identical in a list |
| What does the click do? | focuses the browser window, then switches to the named session if it is not the one on screen | The reference's click target is a tab in its own tab strip. Here one browser tab is one conversation, so the browser window is the tab; app/page.tsx registers that behaviour with registerDesktopNotifyFocusHandler rather than letting the notifier reach into lib/api |
| What does the switch do? | turning it on is the permission request; it latches only on granted |
A switch that reads ON while the browser has refused is a dishonest state, and it would notify nobody. Turning it off never asks anything |
| What is shown after a refusal? | the row's own hint gains a second sentence naming the refusal and the way out, in the status colour | The permission is read live on every mount of the settings page and is never persisted — a stored granted would outlive the user revoking it in the browser's site settings. Reopening the settings page is therefore enough to see the refusal |
| Is the permission ever stored? | no | Notification.permission is owned by the browser and revocable outside the page |
Notification.requestPermission() is only ever called from the switch's
onChange — never at load. A permission prompt raised before the user has
expressed any intent is the pattern browsers penalise, and the switch is
the request. The default is therefore "false": a profile that never touched
the switch asks for nothing and shows nothing.
How to tell it works
- Settings → General → 桌面通知, turn it on, accept the browser prompt.
- Switch to another browser tab (or window) while a turn is running.
- On the settle you get one notification; clicking it brings the window back and lands on that session.
- Raise a turn failure (an unknown
/cmdproduces an[chat.send]anomaly): you get the failure notification, and not a second "finished" one. - Revoke the permission in the browser's site settings and reopen the settings page: the switch reads off and the refusal sentence is there.
The context-window readout
components/context-meter.tsx is the only consumer of
webui-context-window-usage. It reads the key once at mount and then
follows subscribeContextWindowUsage in webapp/lib/settings-local.ts,
so the switch takes effect in the already-open page — the settings modal
and the composer are on screen at the same time, and a reload would be
the only other way to hear about it. The channel's shape is
subscribe*(listener) → unsubscribe, the same one
webapp/lib/theme.ts#subscribeSystemTheme uses for the appearance picker;
a listener that throws is isolated so it cannot cost the other subscribers
their update.
The default stays "false", the desktop reference's own default, so a
profile that has never touched the switch renders no readout. That is a
change from the webui's previous behaviour, where the meter drew
unconditionally and the switch did nothing: the reference hides it by
default, and the switch is what decides. The stored format is still the
bare "true" / "false" string — the key did not move onto the
webui:ui:v1 envelope, which would have broken the reference-shared
contract.
Search and layout details (ticket 48)
- The search matches the localized label and the internal key: typing
custom-instructionsfinds Personalization,usagefinds Usage & models (the nav carries alias entries where this client's tab id differs from the reference's key). - The content column carries the active tab's title in an
<h2>that follows every switch. - Tab switches replay a 180ms horizontal fade-in (
key={active}re-mount); withprefers-reduced-motionthe animation is off and the content renders in place. - The search field is a bordered 36px container — leading search glyph, input, and a clear button that appears only when there is text; the back affordance carries a "Back to app" label.
Present in the reference, not implemented here (recorded honestly)
| Capability | Why |
|---|---|
| Account page | sign-out only — the readings are real (see Account section (账户) below); the engine exposes no sign-in/sign-out method, so the button renders in the reference's form, disabled, with that reason in its tooltip |
| Archived tasks page | the tab renders its empty state; the list, its restore and its delete need the archived-session contract |
| Usage & models three-source switching | the segmented tabs now match the desktop form (ticket 53), but they are a view switcher — they do not switch the model source in use; real Token Plan / MiniMax API / custom-model routing plus source badges still need a model-routing contract |
| MiniMax API key panel | input + connectivity test + save-and-use |
| Custom model drag-reorder, per-model toggles, preset picker | provider contract work; add, edit and delete all converge on one dialog (this batch), leaving the list to display and delete |
| Search keyword highlighting | the reference itself never wired it (component + keyframes defined, no call site) |
| General-page dataDir footer | see the section table above |
Usage & models (用量与模型)
Below the page's h2 sits the desktop's segmented header: 「Token Plan 使用中
⌄」 (selected state is a grey pill; the green "active" badge and the
disclosure chevron render in the desktop's form, and the dropdown itself is
deliberately omitted per ticket 53 — the local edition has no plan source to
switch between) | a hairline | 「自定义模型」. The page lands on the Token
Plan view; the one exception is the model selector's add-provider deep-link
(autoAddProvider), which seeds the custom-models view — otherwise the add
flow would fire behind a view where the panel is not rendered.
The Token Plan view is the desktop's five blocks (the tabs plus four cards):
| Block | Data policy |
|---|---|
| Plan card (ⓘ + two rows + 管理⌄) | The plan NAME is real: tokenPlan.tier from GET /api/account, read when the view mounts, rendered verbatim; when the engine reports no plan, or the account surface is unreachable, the row renders 「未订阅套餐」 — never a default tier. The remaining figures have no credential path here: the credits figure renders 「云端账户域,本网页端无账户凭据」 and the expiry line is omitted rather than given a fabricated date; 升级 (black primary) / 管理 ⌄ / 去充值 render in the desktop's form but disabled, because all three act on the cloud account |
| Usage card (three stacked progress bars) | The 5-hour and weekly windows are the one live source (engine over ACP, POST /api/usage; polled every 2 minutes, manual refresh records a forecast sample): with data they print the desktop forms "X% / 100%" / "X%" plus a relative reset caption ("resets in 43 min"); with no reading a bar shows the unavailable line, never 0%; the video window has no local source and permanently shows 「本地版不适用」 |
| Credits row (ⓘ + blue switch) | Credits are a cloud-account figure: the switch renders the desktop's blue on-form but greyed (checked + disabled), the hint is the reference's, and the row names the cloud account domain as the reason |
| Invoice row | The one fully live affordance: 申请 ↗ opens the MiniMax open platform in a new tab |
The plan card is where ticket 53's A1 ruling was revised. A1 read 「无源即占位」 — a figure with no local source renders the placeholder — and applied that to the whole card, but the local server does have sources here (POST /api/usage for the quota windows, GET /api/account for the plan tier), so the ruling overstated the gap. The revision splits the card by source rather than by card: what the server can read is rendered, and what belongs to the cloud account domain renders the honest line that names that domain as the reason. Two alternatives were rejected — keeping the whole card on the placeholder (a plan name the engine has already reported is not a gap), and wiring credits / expiry / invoicing as well (this session holds no account credentials for the cloud account, so a real-looking figure there is exactly the fabrication A1 exists to prevent).
The custom-models view is the existing provider panel (API keys, protocols, model lists, connection tests, preset one-click enable); ticket 54 rebuilt the add flow into the desktop's dialog form (next section) and this batch folded editing into that same dialog (see "One dialog also edits" below) — the panel is now a list plus a delete affordance, with no second editor.
Account section (账户)
One read of GET /api/account on mount answers the whole section. The
endpoint is not new and not duplicated: it is the same projection the user
menu's account card already reads, fetched on demand because the state
snapshot is broadcast to every SSE subscriber.
| Row | Field | Missing-value sentence |
|---|---|---|
| 账户名 | identity.name, trimmed |
the engine answered and reported no name |
| 当前套餐 | tokenPlan.tier, through the Token Plan card's own planNameOf |
「未订阅套餐」 when the engine reported no plan; 「正在读取当前套餐…」 while a read is in flight; the unread sentence when the surface itself failed |
| 配额概况 | tokenPlanQuotaState, plus quota.fiveHour / quota.weekly remainingPercent |
「引擎未返回读数」 per window; 「不限量」 when the engine reports the window unmetered |
| 账户状态 | status |
the unread sentence; a status token with no dictionary entry resolves to no sentence rather than leaking a raw enum |
The unread sentence is one per failure kind and each names the endpoint:
ok: false renders the engine's own reason when it sent one, a transport
failure renders the read-failed line, and none of them asserts a fact about
the user. The two failures that are NOT the same thing — an unreachable
account surface and an answer with no account name — therefore get different
sentences, which is what the previous hardcoded 「本地模式,未登录」 row
could not express.
Division of labour with the Token Plan card is by shape, not by topic: that
card owns the limit BARS, the plan actions, credits, expiry and invoicing;
this section owns identity and the plain-text readings. The only shared
value is the plan name, and it goes through one resolver (planNameOf) so
the two surfaces cannot drift. A window reported as unmetered prints
「不限量」 rather than 「剩余 0%」 — a plan with no cap must not read as an
exhausted one.
Rejected alternatives: a second account endpoint (the projection already exists and a second route would be a second contract to keep in sync), and showing the quota as bars here as well (the same gauge twice, from two different sources, on two pages).
Add-model dialog (ticket 54, 53b)
The 「+ 添加模型」 button — centered under the empty state, at the bottom of the provider rail once providers exist, and via the model selector's deep-link — opens the desktop's modal instead of appending a rail draft:
| Dialog region | Contract |
|---|---|
| Provider select (「请选择提供商」) | Options are GET /api/providers/presets plus a 「+ 其他(自定义)」 sentinel; choosing a preset fills id / label / auth-type / baseURL and seeds 「API 格式」, choosing the sentinel expands the custom fields (id, display name, auth type); 接口地址 is a top-level field for both branches — see "One dialog also edits" below. DeepSeek / Zhipu AI(智谱)/ Moonshot AI (China) carry the reference's spellings; other local presets keep their catalogue labels. A 404 catalogue degrades to the custom-only dropdown |
| API 格式 | The desktop's second field, rendered for every provider rather than only for 「其他(自定义)」. It is the existing wire protocol under the desktop's labels — OpenAI Completions / Anthropic Messages / Gemini — so no new format reaches the backend. Choosing a preset seeds it from that preset's own protocol and it stays editable afterwards. The protocol select that used to sit inside the custom branch was removed rather than kept alongside: two controls bound to one value is how the preset and custom branches end up disagreeing about what gets saved |
| 自定义 Headers | Rows of (name, value) with 「+ 添加 Header」 and a per-row remove, held as a list rather than an object so a half-typed row survives editing. A blank name is dropped, a name is trimmed but a value is not, and a later duplicate wins — all three decided in one place (headerPairsToRecord), so the dialog, the PUT body and the server cannot disagree. Zero rows render an explicit placeholder rather than collapsing. The collapse result lands in auth.headers on the PUT body and comes back in auth.headers on GET /api/providers |
API key (AntInput.Password) |
The eye toggle is safe here and only here: the field's value is what the user just typed, not a masked placeholder — the editor's no-reveal rule (keep-existing-key convention) is untouched |
| Model entries (「模型 01…」 + connectivity test + ↻ reset + 🗑 delete) | Five fields: name → id, context window → contextLimit, max output tokens → disabled with the 「本地版不适用」 marker (the /api/providers PUT contract has no field to persist it), reasoning levels → thinkingLevels fed from THINKING_LEVELS (the low/medium/high contract is frozen; the reference's 「max」 placeholder example is deliberately not copied), attachments → four checkboxes 图片/PDF/视频/音频 mapping to image/file/video/audio (file joined MODALITIES; text passes through untouched) |
| 「+ 添加」 / 「自动获取」 | Add appends a blank entry; auto-fetch opens the 「已获取模型」 dialog listing the selected preset's built-in catalogue with the note that it is not a live per-key query — the local backend has no model-listing proxy. With no preset selected the dialog states the missing capability instead of inventing rows. 「全选(n/N)」 + 取消/添加 follow the reference; picked entries arrive with their catalogue metadata. Both actions carry tooltips spelling out the split (manual entry vs catalogue pick; auto-fetch reads the list only and saves nothing); with zero entries the models section renders a dashed placeholder naming both paths instead of collapsing |
| 跳过连通检测 / 连通检测 | The desktop's form-level check, on the left of the footer bar. It reuses the existing POST /api/providers/test contract with the live form values (protocol, key, baseURL and the custom headers, so the probe exercises the request that will actually be sent) and records one verdict, rendered as 「可达 · Nms」 / 「不可达:错误」. This is a different scope from the per-entry 检测 on each model card — that one asks whether a model id responds, this one asks whether the provider is reachable at all — so both exist. Editing any probed input (provider, API 格式, key, baseURL, auth type, any header row) drops the verdict, because a verdict that survived an edit is a pass for a request the provider will never see. Header values are re-validated server-side on this path too, since the test endpoint takes auth straight from the body without the PUT normaliser |
| 取消 / 保存 | Save validates (provider chosen, unique id, per-entry validateModelRow), appends the draft to the panel's list, and PUTs through the unchanged draftToWire + api.putProviders({version: 2}) path; on failure the dialog stays open with the typed input intact. The pair sits in a dedicated footer region (hairline separator + 16px clearance) at the h-9 control height, the black primary carrying the token shadow. 保存 is disabled until the form-level check passes, matching the reference footer and the greyed 保存 in the reference screenshot; 跳过连通检测 is the escape hatch for an operator who cannot reach the endpoint, and a line states which of the two is blocking. A disabled control with a stated reason and two ways to lift it is not a dead button |
Ticket 54 invariants — no server-contract change (the /api/providers PUT
body, /api/set-model, and every endpoint are untouched; the whole delta is
client-side plus tests and docs); every pre-existing data-testid on the
panel/editor surfaces survived in source at the time
(webapp/test/add-model-dialog.test.ts pinned 36 + 2) — that assertion is
now void: after the flat editor's retirement its testids are pinned as
must-not-return instead, while the list chrome's testids are still pinned
one by one; the preset catalogue, thinkingLevels editing semantics,
provider grouping and the thinking-display exceptions are unchanged; the
auto-add deep-link still lands on the custom-models view, now opening the
dialog. The legacy rail-draft addProvider path and the editor's dead
auto-focus prop were deleted with their behaviour subsumed by the dialog.
Acceptance round 2 (same ticket). The dialog components moved to
components/add-model-dialog.tsx and export their controlled surfaces,
which unlocked behaviour-level tests (renderToStaticMarkup, the 53a
F-7 harness): the eye round-trip, the validation error block, the
cancel-reset landing state, the check-all semantics with its n/N
counter, and the zero-selection / custom-provider disabled states are
now pinned on rendered markup plus the pure helpers — an 11-mutation
spot-check that reverted dialog behaviour all fails red. The PUT-body
red line moved from a call-site literal to a closed key-set assertion
on draftToWire itself (provider-management.test.ts). One wording
correction from the first round: with a custom provider the 自动获取
link is NOT disabled — it opens the dialog, which states the missing
capability honestly and disables 添加. The count correction: 36
preserved testid literals in the panel-source pin plus the 2
empty-state affordances = 38, matching the base tree.
Ticket 56 — dialog visual/interaction parity with the official
screenshot (user report: ugly, interaction/layout inconsistent;
reference design-ref/screenshots/byok-custom-model-official.png):
- Layout. Both modals centre vertically (antd
centered); the card takes--radius_12with an elevation composed from the--opacity_black_1_8/1_15ramp (no literal rgba). 取消/保存 moved into a dedicated footer region (hairlineborder_default+ 16px clearance), buttons at the h-9 control height with the black primary carrying--shadow_default. The models header packs its label and actions adjacently — the oldjustify-betweenrow read as broken through its wide dead gap. The form body clamps to90vhwith an internal scroll and a footer that never scrolls away — the live verify round caught the filled custom branch (5 provider fields + entry cards) reaching 884px in a 633px viewport, where antd's overlay offers no scroll and 取消/保存 sat below an unreachable fold. - Models empty state. With no entries the section renders a dashed placeholder naming both add paths instead of collapsing to blank; +添加 / 自动获取 carry tooltips stating their division of labour.
- Connectivity test (per-entry 检测 button). The official
semantics is "probe with the currently filled-in info, per model".
The local implementation reuses the server's existing
POST /api/providers/testcontract verbatim — protocol whitelist, local key-format check, then a real fetch against the configured baseURL — with no new route. The dialog shell assembles the probe from the current form values (the protocol comes from 「API 格式」, the endpoint from the top-level 接口地址 field — which this batch lifted out of the custom branch, so a preset is no longer pinned to its catalogue endpoint) with a 4s timeout, and renders 「可达 · Nms」 (success token) / 「不可达: error」 (error token). Granularity, stated honestly: the probe is endpoint-level (baseURL + key) and does not exercise the entry's model id — the tooltip and this paragraph say so rather than claiming the reference's model-level coverage. The button's availability mirrors the server's local gate: a byok probe demands a typed key, a coding-plan probe (the claude-code / codex / opencode-go presets) fires with the endpoint alone. A verdict is dropped the moment its inputs go stale — editing/resetting an entry drops that entry's verdict, removing one shifts the rest, and any change to the shared probe inputs (provider choice, protocol, baseURL, auth type, API key) drops them all. The ticket also fixed the probe's target:testProviderpreviously read and validated the body'sbaseURLbut dropped it, sending every probe to the protocol default; it now honours it as the route's comment always claimed (PUT /api/providersis untouched). - Auto-fetch semantics vs the official table ("read the provider's model list for selection; saves nothing"). The local behaviour already matches — picked rows only land in the dialog draft, and saving is a separate explicit action. The one difference is the list's source (the local preset catalogue, not a live per-key query), which the fetch dialog already states honestly; no behaviour change was needed.
Ticket 85 — the three desktop fields the local dialog was missing
(API 格式, 自定义 Headers, footer 连通检测 / 跳过连通检测)
The add surface was already dialog-based on this branch — that part of the request was a no-op, and no work was spent re-doing it. What was genuinely missing is below; the per-model-entry 检测 from ticket 56 and the page-inline editing flow are both untouched.
自定义 Headers is the only one of the three that needed a contract
change, so it is the one worth reading closely:
| Layer | What it does |
|---|---|
| Dialog | Rows of (name, value). Held as a list, collapsed by headerPairsToRecord (blank names dropped, names trimmed, values not, later duplicate wins) |
PUT /api/providers |
New optional field providers[].auth.headers: Record<string,string>. A body that omits it is byte-identical to the pre-ticket body and a stored provider without it loads to {} — the field is additive, not a migration |
GET /api/providers |
Returns auth.headers verbatim, unmasked |
engine-provider-sync |
Copies it to the engine's options.headers, omitted when empty. This is the load-bearing link: local-runtime-v2 already merges options.headers into every upstream request for the provider (catalog/provider-views.ts:218), so no runtime change was needed |
POST /api/providers/test |
Carries the headers into the probe, re-validated through the same grammar because this route takes auth from the body without the PUT normaliser |
Why headers are not masked, when apiKey is. apiKey is masked
because the server substitutes it on the operator's behalf — the operator
never needs to read it back. A custom header is routing or tenant
configuration the operator typed and must be able to edit, so masking it
would create a write-only field. An operator who treats a header value
as a secret has no way to express that here; the honest statement is that
this field is not a place to keep credentials. The API Key field remains
the only masked one.
Validation is a reject, not a strip. Names must match the RFC 9110 token grammar and values may not carry CR, LF or NUL; a record that fails rejects the whole PUT with an error naming the provider and the header. Silently stripping the character would leave the operator believing a header is in effect when the upstream never received it intact. Ceilings: 20 headers, 128-char names, 4096-char values.
Probe asymmetry, stated. The connectivity probe spreads operator
headers FIRST, so the protocol's own required headers (Content-Type,
anthropic-version, Accept) overwrite them. A probe answers "can I
reach this provider", not "replay my headers exactly"; letting a
mistyped Content-Type break the probe would make it answer a question
the operator did not ask. The production request path has no such
restriction.
Still not done (deliberately, for the next batch): the desktop's 「模型 01」 nested sub-card with 模型名称 / 上下文窗口 / 最大输出 Token is a screenshot-only shape — the reference implementation carries a single 模型名称 textarea instead, so the desktop is newer than the reference and there is no second source to check it against. Rebuilding the model-entry structure is a larger change than this batch and is left alone.
This batch — one dialog also edits, the flat editor is retired, and 「API 格式」 drives the endpoint surface.
| Change | Contract |
|---|---|
| Editing runs through the same dialog | A click on a list row opens the same modal (editTarget), seeded by editSeedFromDraft: a preset provider lands on its own catalogue branch, a custom one on 「+ 其他(自定义)」. Id, display name, auth type, 接口地址, the custom headers and the model entries all arrive pre-filled. The title switches from 添加模型 to 编辑模型 (providers.dialog.editTitle). Two entry points share one form rather than two forms free to drift |
| The commit starts from the stored record | The committed draft begins at the record being edited (const base = editTarget ?? newDraftProvider()); only the fields the form can reach are taken from form state. enabled, preset and draftId are properties of the record, not the form — rebuilding one would re-enable a provider the operator had disabled, or detach it from its preset. The panel locates and replaces the row by draftId, not by wire id: matching on the id would write a second record instead of renaming the first whenever a custom id is retyped |
| The key still never lands on disk in edit mode | The edit seed's API Key is always "", the server's keep-the-existing-key sentinel; the masked value reaches the placeholder only and is never written back as a value. An untouched field preserves the stored credential; a typed one replaces it |
| The flat editor is retired | ProviderEditor and its ApiKeyInput / DraftModelList / DraftModelRow, plus the panel-level 「保存供应商」 button, the whole-list validation and the selected-row probe, are deleted. Delete did not disappear with them: it moved onto the list row as 🗑 (provider-delete-{draftId}, Popconfirm confirmation). Preset rows still carry no delete, the retired editor's own rule — a preset's lifecycle belongs to the preset controls |
| The row element changed | The row is now a <div> wrapper rather than a <button>: it holds a second clickable control, and a button inside a button is invalid HTML with ambiguous keyboard semantics. The row's edit control is provider-row-edit-{draftId}, its delete is provider-delete-{draftId} |
| 接口地址 lifted to a top-level field | The field used to sit inside the 「+ 其他(自定义)」 branch, which meant 11 of the 12 preset providers could neither see nor change their endpoint. It now renders for every provider: a preset choice seeds it with the catalogue endpoint, the operator may override it, and an empty field falls back to the same protocol default the server uses |
What the 「API 格式」 selection drives (API_FORMAT_SPECS, every value
transcribed verbatim from server/lib/providers-config.js —
DEFAULT_BASE_URL and probe()):
| Format | Default endpoint | What 连通检测 actually sends | How the key travels |
|---|---|---|---|
OpenAI Completions |
https://api.openai.com |
GET {baseURL}/v1/models |
Authorization: Bearer header |
Anthropic Messages |
https://api.anthropic.com |
POST {baseURL}/v1/messages |
x-api-key header |
Gemini |
https://generativelanguage.googleapis.com |
GET {baseURL}/v1beta/models?key=… |
?key= query parameter, not a header |
The hint under the field shows the resolved probe target: the typed endpoint when there is one, the format's default when the field is blank — i.e. the request the test button will send. The credential line switches with it, and Gemini's says outright that the key rides in the query string; copying the other two formats' "key in a header" habit yields a 401.
KNOWN DEBT — per-format field show/hide was not built, and here is
why. The three protocols carry the same field set in the
/api/providers PUT contract; the backend has no field to hide per format.
The linkage therefore lands on dynamic content (the endpoint default, the
probe request, the credential's transport) rather than on a field's presence:
hiding a field this backend cannot store would produce a form that lies about
what it saves, which is worse than not hiding it. Real per-format show/hide
requires extending the PUT contract and is its own batch.
This batch's invariants. No server-contract change (all five
/api/providers endpoints and /api/providers/test are reused as they are;
the whole delta is client-side plus tests and docs). The PUT body is
byte-identical — still draftToWire + api.putProviders({version: 2}), with
an edit substituting one record in the list. enabled, preset, the masking
convention, the preset catalogue and the thinkingLevels semantics are all
unchanged, and the engine layer, sessions.js and markdown were not touched.
On the test side the list chrome's testids stay pinned positively, the
retired editor's eleven are pinned negatively (a return turns the suite red),
two named suites cover the format linkage and the edit reuse, and ten revert
mutations were run against them — all ten were caught.
Ticket 53 invariants — no server-contract change (the delta:
panels.tsx / usage-models-cards.tsx / icons.tsx / i18n.ts plus two
test files); the h2 header, the per-tab fade-in and the 760px page width are
untouched; SETTINGS_NAV, the SettingsSection union and the deep-link
entry points (initialSection, autoAddProvider) are unchanged; the eight
tabs that were placeholders when this round landed kept their placeholder
form — the settings-modal port (58) and its four sub-pages (55a) later gave
most of them content, see the navigation table above. The usage.used /
usage.reset label strings, which lost their last consumer to the
desktop-figure forms, were deleted from both dictionaries.
Acceptance fixes (second round, 2026-09-29). The usage-bar track moved
to bg-border_default — the context meter's bg-bg_grouped_tertiary_elevated
it had been copied from resolves to the same grey as the section card in
the light theme (--gray_75 both), making all three bars invisible, and
the reference itself is light-themed. The invoice 申请 ↗ action switched to
the white-ground grey-outline form the 去充值 / 管理 buttons share. Whole-hour
reset captions drop the minute slot ("resets in 1 h", new key
usage.duration.hour). The four pure display cards now live in
components/usage-models-cards.tsx, and
webapp/test/usage-models-cards.test.ts asserts their RENDERED markup
(renderToStaticMarkup: placeholders, disabled actions, the track token,
the outbound link, the caption boundary) — the track-token assertion was
red-green verified: injecting the old token fails the test.
The user menu's usage row
The row used to host a hover flyout with the quota figures. It now jumps
straight to the settings page's Usage & models section (onOpenUsage →
initialSection: "providers"); the flyout component and its
usagePopover.* strings were deleted. There is exactly one surface for
quota figures.
The nav item id behind the section is "providers", unchanged: the model
selector's "Add provider" deep-link targets that id, and only the visible
label moved.
Ticket 48 invariants — what this round did NOT change: the server
contracts (server/routes/settings.js, server/routes/providers.js,
server/lib/settings.js are untouched — the whole delta is client-side);
the SETTINGS_NAV four-group division and the three-value
SettingsSection union; the deep-link entry points (initialSection,
autoAddProvider — the model selector's add-provider flow and the user
menu's usage row both still land where they did); and the eight tabs that
were placeholders in this round, which the settings-modal port (58) has
since given content. The dead if (!section) branch inside SettingsPanel
was removed and the section prop made required — every reachable tab
resolves a section, so the branch could never render.
What the user sees. Opening the 用量与模型 tab reads the engine
once and settles three things that used to be local guesses: which
credential the engine is actually using, whether a MiniMax API key is
stored, and what the last connectivity probe found. The pill is still
the view; the 「使用中」 badge beside it is the engine's answer, and it
moves only when a write has been confirmed. Picking Token Plan or MiniMax
API in the dropdown switches the view AND writes the engine
(PUT /api/model-source); a refusal — the engine's NO_API_KEY when no
BYOK key is stored — leaves the view where the user put it, so the key
field they need is the panel that stays on screen, while the badge keeps
showing what is really in use.
Why the badge and the view are separate values. They were one value
before, which is what made the old build's claim false: a useState
switcher could render a source as selected while the engine kept using
the other one. A badge that claims 使用中 for a source the engine never
accepted is the fake-success shape this codebase keeps refusing, so the
badge is fed exclusively by a read-back.
The key row. A stored key shows as the engine's mask, never as
plaintext, and typing a new value replaces it on save. 保存并使用 is one
request, not two: the engine writes the key and switches the source in a
single transaction, so the tab never shows a saved key beside a source
that was not switched. An empty submission is the keep sentinel
(changed: false, no engine write) — the same convention
PUT /api/providers uses, and it exists because the read can only
return a mask while the engine rejects a mask submitted as a key.
What the probe does and does not test. 检测 probes the STORED key on
the minimax_api provider and the response says so (tested: "stored_key"). Two limits are the engine's contract, not this UI's:
testUserModel takes no key override, so an unsaved value cannot be
probed — the button is disabled while the field holds one, and a visible
line under the field says why, because the reason used to live only in a
title attribute that keyboard and touch users never see — and the
managed Token Plan credential is not a model-service key, so the Token
Plan source has nothing to probe here. A probe that ran and failed is a
completed probe, not an error: it renders the engine's status.
The key row is four states, not two. The badge reads the engine's
masked projection, but it is describing a field the user can be editing
right now, so a typed-but-unsaved key is its own state (「已输入,未保存」)
that outranks both 「已保存密钥」 and 「未启用」 — the user is replacing the
stored key, or has plainly typed one, and neither badge is true. Symmetric
to that, the engine's NO_API_KEY refusal is a verdict on an EMPTY field:
typing one falsifies it, so the pinned 「请先填写 API Key」 is dropped on
the next keystroke. A refusal that is not about the missing key — a
transport failure, a rejected write — is still true afterwards and stays
on screen.
Cost. One extra read per settings-tab open. The read boots the
engine runtime if none is up, which is the write-side contract and is
acceptable here because the user opened the tab; a future change that
moves this fetch to page level must use the non-booting host getter
instead (server/engine/model-source.js KNOWN DEBT 2).
What this did not do. The add-model dialog's 「自动获取」 still
resolves against the built-in preset directory: v2 has no per-provider
catalogue query for an arbitrary key, so a live per-key fetch has no
engine method behind it. The Token Plan cards stay on decision A1
(本地版不适用) — wiring them to /api/usage and /api/account is a
separate, undecided item, not a side effect of this one.
The defect this fixes. A UAT round trip on 2026-10-03 (板块 4) switched
Token Plan → MiniMax API → Token Plan with a key stored in between. The
Token Plan card came back reading 「未订阅套餐」 while GET /api/account
answered tier: "Ultra" throughout, and only F5 recovered it. The
endpoint was never wrong.
Root cause. The plan name is read in UsageModelsSection
(webapp/components/panels.tsx), mounted only while the port's view is on
the token-plan tab, so a switch away and back remounts it. The read was
useEffect(..., []) — once per mount — and that mount raced the engine:
the section rendered in the same tick as PUT /api/model-source, and
GET /api/account answers an engine that is still rebinding with HTTP
200 and {ok: false, reason: "no_client"}. The card read that as "no
plan", nothing re-read it, and the section's state outlived the failure.
The re-read. The port owns an accountRevision counter and increments
it after every successful source write — the dropdown's PUT /api/model-source and 保存并使用's PUT /api/model-source/api-key, which
switches the source inside the same engine transaction. The section's
/api/account effect lists that counter as a dependency, so a confirmed
write re-runs the read against an engine that has finished rebinding. A
refused write bumps nothing: nothing changed, and a re-read would only
spend a request to re-render the same answer.
A failed read may not un-know a name. Revalidation alone is not enough,
because the re-read can also lose the race. reconciledAccount
(webapp/components/usage-models-cards.tsx, pure and unit-tested) keeps
the last ok: true answer standing: only an ok: true payload is new
information, so an unreachable account surface cannot knock a known plan
off the card. An ok: true answer that reports no plan does replace it —
the engine saying "no plan" is an answer, saying "unreachable" is not.
A read in flight is its own sentence. With the name held, the only remaining nameless state is "not read yet", and it renders 「正在读取当前 套餐…」 rather than 「未订阅套餐」. An account surface that has not answered has not said the user has no plan — that conflation was the visible half of UAT4-1.
Cost. One GET /api/account per confirmed source write, on a surface
the user has just acted on. The read is not debounced: a source switch is
a deliberate act, not a stream of them.
What the user sees. 跟进消息行为 has three positions. While a task is running, the composer keeps Stop where it was and — with 排队 or 立即发送 — also offers the send arrow: 排队 hands the message to the engine's queue so it runs after the current turn, 立即发送 steers the running turn. With 关闭 the composer is exactly what it was before: the send control is replaced by Stop until the turn ends and the text waits in the box. Flipping the switch takes effect in the open page; no reload, no new send.
Why there is a third position. The desktop reference's control has two, because the desktop owns the running turn and neither of its options can fail. webui's follow-up can be refused — the engine that owns the running turn may be another process — so a two-valued switch would be a behaviour change wearing a switch's clothes. 关闭 is webui's own option and is documented as such.
The ownership gate, and why it is not a transport check. Before
either action runs, the server asks the engine whether this process owns
the active turn (cliService.getActiveTurn). Under the runtime
transport the turn runs in the webui process, so the answer is yes and the
queue or the steering message is admitted. Under the default acp
transport the turn runs in an mcode acp subprocess, and queueing into
this host would wake its own dispatcher and start a SECOND turn for a
session that already has one — the failure /api/send spends four claims
and a 409 preventing. That case is refused with turn_not_owned, the
text comes back to the box, and the banner says why. The gate reads the
engine's own answer rather than MCODE_WEBUI_TRANSPORT, so it stays
correct when the chat path finishes its move to the in-process transport
and needs no edit to start working.
What the response reports. The engine's own answer, never the request: a queue answers with the item id and position the engine committed, a steer with the turn id and delivery mode. The two refususes stay two different sentences, because "no turn any more" and "the turn is in another process" call for different next steps.
Cost. One engine read per follow-up send, and a write that boots the runtime if none is up — acceptable, because the user pressed send.
What this did not do. The queue has no UI: a queued message is committed with an id and a position that nothing displays, and cannot be inspected, reordered or cancelled from the browser. That is PB-13's scope, and until it lands a queued follow-up is invisible until the running turn ends. A steered message reports admission, not whether the running agent read the text before its next step.
What the user sees. The About section's first row is 导出日志 / Export
logs, and its button works: the browser saves one text file named
mcode-webui-logs-<timestamp>.txt. The row's description says what is in
it — this server's error log and recent activity — and adds that nothing is
uploaded anywhere.
Why the rename. The row used to be a permanently disabled button labelled 上传日志 / Upload logs. There is no upload service in this edition: no telemetry sink, no ticket intake, nothing that leaves the machine. A label that names a destination the product does not have is a promise, and a disabled control cannot keep it — the tooltip only contradicted the title. The alternative considered and rejected was keeping 上传日志 and pointing it at the download, on the argument that the desktop reference uses the word; a reference's word does not make its destination real, and the button would then say "upload" while writing to the user's disk.
What the file contains. Two sources, both read from the module that writes them, so a relocated data directory cannot make the export silently empty:
| Section | Source | Bounded by |
|---|---|---|
| Server error log | WEBUI_DATA_DIR/.server.err (config.js#installGlobalErrorHandlers) |
2000 lines / 2 MiB |
| Event log | events.path() (honours MCODE_WEBUI_EVENTS_PATH) |
2000 lines / 2 MiB |
The tail, not the head: the failure being investigated happened most
recently, and the event log is tens of megabytes on a long-lived install.
Every bound is printed in the file itself ([truncated: showing the last N of M lines]), so a reader can tell a bounded file from a complete one
without trusting the tool that produced it.
What the file deliberately does not contain. sessions.json
(conversation transcripts), settings.json (provider credentials) and
uploads/. A diagnostics file users attach to a bug report must not be
the one file on the machine carrying their API keys and their
conversations. The source list is a closed two for that reason, and a test
asserts the markers never appear in a bundle built next to decoy files.
Why there is no failure status. The endpoint answers 200 in every
case, including a log file that does not exist. The client is a browser
anchor with a download attribute, so a 404 would be saved into the
user's downloads folder as mcode-webui-logs-<timestamp>.txt containing a
JSON error body — a file that looks like logs and is not. Absence is
therefore reported where the reader is: inside the body, per section.
What it costs. One read of two local files per click, synchronously on the server's event loop. Both are capped, so the worst case is a few MiB of already-warm page cache.
Known debt. The bundle is plain text with no redaction: an engine error line can quote a prompt fragment. Redaction is not attempted because there is no reliable rule for what is secret in an arbitrary log line, and a partial redaction would be worse than none. Until the engine's own logging grows a redaction hook, the operator is the one deciding what to share from the downloaded file.
What the user sees. 自定义指令, 关于你 and 自定义审查准则 each keep their textarea and their 保存 action, and each now prints one line under the field: 「已保存于本浏览器,不会注入引擎会话。」 / "Saved in this browser only — it is not injected into engine sessions."
Why. The texts have always persisted — that part is real and stays
real. What is not real is any claim that they reach the engine. A grep of
the runtime source for setConfigOption — the option the plan had assumed
would carry them — found nothing: the method does not exist anywhere in
local-runtime-v2, so there is no config write channel to hang them on,
and no session-creation parameter that takes them either. With no consumer,
a saved 「自定义指令」 read as an instruction the agent follows. The field
was making a capability claim its backend had already disproved.
The decision tree this took. The three options were: extend the engine
contract, drop the fields, or keep the storage and stop claiming the
effect. Extending the contract is engine work with no local caller to size
it against; dropping the fields removes a place users keep text they own
and can read back at any time. What remains is honest and cheap: the
storage is a user's own local text, the field says plainly that it is not
an instruction, and the moment the engine grows a channel, deleting one
sentence and one note prop is the whole change.
What this does not do. It does not make the texts work, and it does not
hide that they do not work — that is what the note is for. Nothing reads
the three localStorage keys: a future change that wires one of them must
remove the note in the same commit, or the field will be describing an
effect it no longer lacks.
The user asked for every desktop main-surface screenshot to be copied verbatim. This ticket covers three elements under ticket 53's A1 ruling: what has a local data source is real; what does not renders the desktop's exact shape with the 本地版不适用 marker, never faked data. Server contracts are untouched — the whole delta is client-side.
User menu (sidebar footer avatar, ref-01) now carries the desktop's
full row set — Settings (with a Ctrl+, kbd badge; the binding is real,
added to app/page.tsx's keydown handler) / Upgrade / Daily check-in /
Usage / Feedback & help / Sign out — plus a trailing user card (avatar,
display name, plan badge, bell). The enabled/placeholder split:
| Row | State | Reason |
|---|---|---|
| Settings | enabled | opens the existing settings modal; Ctrl+, is a new real binding |
| Usage | enabled | jumps to the settings page's Usage & models section (2026-09-28 decision, unchanged) |
| Upgrade / Feedback & help | disabled placeholder | cloud billing and product support pages; hover title carries common.notLocal |
| Daily check-in / Sign out | disabled placeholder | engine contract not landed (pre-55c treatment kept); hover title carries common.unsupported |
| User card | real | engine identity and plan badge when /api/account reports them; the 本地用户 stand-in and no badge otherwise. The bell opens the existing inbox flyout with the unread dot |
The desktop's UID line at the top is deliberately not rendered: a local edition has no account id to print, and an empty or invented id would break the honesty rule the rest of the menu follows. This reverses the 2026-09-23/24 trims of the Upgrade / Feedback rows — the earlier call held that a disabled row next to a real one was the shape the user had rejected; the 2026-09-29 instruction to copy the screenshots overrides it, with the A1 marker keeping the limits stated rather than implied.
Project context menu (right-click a sidebar project row, ref-26): 重命名项目 / 置顶项目 / 在文件夹中显示 / 归档对话 / 移除 (danger).
- Rename and pin are real, backed by a browser-local overlay
(
webui:project-custom:v1, see Persistence keys). mcode's runtime db has no project entity to write into — a project is the git root its directories resolve to (server/lib/session-tree.js#buildTree) — so the overlay lives where its only consumer lives, mirroring howtitleCustomoverlays session titles. Pinned projects sort to the top and carry a persistent pin mark beside the title; the menu row toggles between 置顶项目 / 取消置顶. - 在文件夹中显示 is live (SB-6). It posts the project's path to
POST /api/fs/reveal— the endpoint has been implemented and registered all along (server/routes/fs.js#handleFsReveal,server/app.js); the menu row was a placeholder that claimed a browser cannot reach the OS file manager, which is false for a webui install (the server holds the workspace). The row is disabled for exactly one reason: the project is bound to no local directory, and the tooltip says so in those words rather than shrugging with 「本地版不适用」. A reveal that succeeds is silent — the file-manager window is the feedback, and a toast would race it. A failure reports through the same banner as the menu's other writes, labelled with the menu's own localized name, whether the server refused with a structuredcodeor the request threw. - The SESSION-level 在文件夹中显示 stays a disabled placeholder. The desktop reference disables it too, so there is no parity to chase and no local limitation to blame — unlocking it would be a product decision this build has not made.
- 归档对话 is a disabled placeholder: the runtime db has an
archivedflag, but writing another process's database is out of scope, and until ticket 55b's archived-tasks page lands there is no un-archive surface — archiving would be irreversible data loss. - 移除 is a real danger row: the confirm modal states the TRUE deletion
set — main sessions AND subagent rows, not the sidebar pill's
main-session count; an irreversible confirm must not understate what
goes — together with the irreversibility and, because the server
authorizes each single-session delete separately through
authorize("session.delete")(there is no batch contract), the number of approval prompts to expect. Once confirmed, the dialog shows livedeleting i/Nprogress while the delete walks the existingDELETE /api/sessions/:idsequentially over every session under the project including subagent rows (the runtime rows are per-session; DELETE does not cascade into children). A failure reports through the action-error banner and stops the batch; the project's rename/pin entries are cleared ONLY on a full success — a partially-failed remove leaves the surviving project its customizations — and the clear runs through the tree's state so memory and localStorage stay in step.
Home quick-capability capsules (ref-28): the home screen renders the desktop's five chips under the composer row — 视频生成 (H3 badge) / Vibe Coding / 设计视觉 / 产品运营 / 询问 MCode — as white pills with hairline borders and the H3 badge in its fixed brand tint. The skills are cloud-only, so every chip answers a click with the 本地版不适用 toast and sends nothing: same shape as the desktop, limit stated in one sentence, no launch to fake.
The regression pins live in webapp/test/shell-elements-parity.test.ts
in three layers: bilingual coverage for every 55c key with the zh labels
asserted verbatim against the reference screenshots; static-source
tripwires for the row sets, the enabled/disabled split, the danger tone,
the batch-delete wiring (including that the confirm quotes the TRUE
deletion set, never the pill's main-session count) and the capsules'
click-to-toast wiring (including a no-api.send assertion on the capsule
strip); and BEHAVIOUR tests on lib/cap-toast.ts — the toast state
machine is an import-clean module precisely so the click→replace→dismiss
contract runs under node:test without a render harness (QA round M6:
hollowing the handler body had kept every source-only assertion green).
Each layer was red-green verified against its own targeted mutation.
Assistant messages and Markdown file previews render through
webapp/lib/markdown.ts (marked, already a workspace dependency — no
CDN). A fenced code block whose language token is mermaid renders as a
diagram instead of a code block. The behaviour is a contract, not an
implementation accident:
| Aspect | Contract | Backed by |
|---|---|---|
| Fence language | The bare token after the fence opener must be mermaid, case-insensitive; trailing metadata (```mermaid {theme: dark}) still matches |
webapp/lib/markdown.ts:133-136 |
| Renderer seam | Fence languages dispatch through a language→renderer registry; the markdown main flow never branches on a language name. Any other fence language can be taken over the same way — one registerLanguageRenderer(...) call. That is the seam a future minimax-code-plugin renderer will install through |
webapp/lib/markdown.ts:54-113, webapp/lib/mermaid-renderer.ts:64-71 |
| Theme | Diagrams re-render when the app switches light/dark: the host watches <html>'s class and mermaid is re-initialised per theme |
components/markdown-html.tsx:52-66, components/mermaid-block.tsx:156-166 + 223-229 |
| Failure state | A syntax error does not blank the page. The failing diagram shows its error text plus the original source in a copyable <pre> — the copy round-trips byte-exact, including `--> |
label |
| Sizing | Diagrams scale to the column width; a diagram wider than its card scrolls inside the card | webapp/styles/mermaid.css:62-80 |
| CJK labels | Node and edge labels render through a font stack with PingFang SC / Microsoft YaHei / Noto Sans CJK SC fallbacks, so Chinese text does not come out as tofu | components/mermaid-block.tsx:109 |
| Loading | The chart library is several megabytes and is import()-ed when the first diagram of a page mounts; the chunk ships with a one-year immutable cache, so later page loads fetch it from the browser cache. A page with no mermaid fence never requests the chunk |
components/mermaid-block.tsx:54-66, server/lib/static.js:64-66 |
| Outline | A diagram is never a heading: the fence emits a <pre>/<div> placeholder pair, not h1–h6, so diagrams appear in no heading-derived outline (the outline the preview renders since slice 27 walks the rendered DOM's h1–h6, which the placeholder pair never produces) |
webapp/lib/mermaid-renderer.ts:44-57, webapp/lib/markdown-toc.ts |
The mermaid dependency (11.12.1, MIT) is recorded in
release/dependency-licenses.json.
HTML written into a Markdown source — an assistant message, an activity
group, a .md file preview — is shown as the source text. It is never
parsed into DOM elements. That is the whole contract; the rest is how it is
kept, and how you would notice a regression.
| Aspect | Contract | Backed by |
|---|---|---|
| Author HTML | Inline and block HTML are escaped, so a pasted <svg><path …/></svg>, a <div onclick=…>, or a <script> block appears verbatim as text. The prose around it is unaffected — one snippet does not blank the message |
webapp/lib/markdown.ts (the renderer.html override) |
| Why the parser and not the sanitiser | marked has no "no raw HTML" option: without the override the snippet reached the tag allowlist, which admits svg/path for KaTeX geometry, and components/markdown-html.tsx then called createElement("path") — an unknown host element, and one The tag <path> is unrecognized in this browser console error per occurrence (nine in a single UAT round; doc/uat/2026-10-03-16-master-sub-agent-comm-redline1.md, anomaly #2) |
webapp/lib/markdown.ts, components/markdown-html.tsx (htmlToReact) |
| Generated HTML is exempt | Markup this app generates never passes through that override: KaTeX arrives from the webuiMath inline extension and every fenced language from registerLanguageRenderer, both of which return their HTML directly. Formula geometry — real <svg>/<path> — therefore still renders |
webapp/lib/math-renderer.ts, webapp/lib/markdown.ts (safeLanguageRenderer) |
| Line structure | A block snippet keeps its original line breaks; escaping never collapses a multi-line paste onto one line | webapp/lib/markdown.ts, webapp/test/markdown-raw-html.test.ts |
| How to tell it works | The regression suite asserts the React tree, not just the string: no svg/g/path element is ever created for author HTML, and a formula still creates svg + path. A change that lets a tag through turns the suite red, and so does one that over-tightens and kills formula geometry |
webapp/test/markdown-raw-html.test.ts |
Rejected alternatives: rendering model-authored SVG (an XSS surface — an
<svg> can carry <foreignObject>, animation and event handlers, and the
product intent is a transcript, not a renderer); adding DOMPurify (a
multi-megabyte dependency to defend markup the app never needs, when the
parser can refuse it outright).
Every markdown codeblock — chat messages, activity groups and markdown
file previews alike — renders through one host
(components/markdown-html.tsx) with the shell the parser emits
(lib/markdown.ts: .codeblock-shell > .codeblock-toolbar +
pre.codeblock-pre > code.codeblock-code; the scroll container is
the code element). The contracts:
| Aspect | Contract | Backed by |
|---|---|---|
| Wrapping | The file_line_wrap switch (ticket 48's key, no new key) extends to markdown codeblocks: on, code lines wrap at the column edge (white-space: pre-wrap; overflow-wrap: anywhere) and the horizontal scrollbar is suppressed; off (scroll mode), lines stay on one row. The language label sits in the toolbar outside the scroll container and never wraps. Read once per host mount — same semantics as ticket 48's file previews: blocks mounted after the toggle reflow, the ones on screen do not |
components/markdown-html.tsx, webapp/styles/markdown-overrides.css |
| Scrollbar visibility | In scroll mode the idle scrollbar is visible: faint grey thumb (8 % opacity token, theme-flipped) over a transparent track, deepening to --utility_scrollbar (15 %) on hover — upstream's sheet painted the idle thumb fully transparent and collapsed the chat-content webkit bar to height:0, so users read clipped code without knowing a bar existed |
webapp/styles/markdown-overrides.css |
Scroll container is a blockified <code> |
The parser emits a bare inline <code> (no .shiki wrapper, unlike upstream markup), and overflow is ignored on inline boxes — upstream's overflow:auto on the element therefore never produced a scroll container here, which is the deeper half of the "can't scroll, can't see a bar" report. The override sheet blockifies it (display: block) so the upstream scroll declaration takes effect; without that line every scrollbar rule is dead styling. A tripwire test also rejects any bare code/pre selector in the sheet, because one would restyle code.inline-code (inline code in prose) |
webapp/styles/markdown-overrides.css, webapp/test/markdown-code-wrap.test.ts |
| Overflow containment for a codeblock taller than the 45vh shell | .codeblock-shell caps itself at max-height: 45vh, and its flex children are the toolbar and the <pre>. The <pre> is not the scroll box — the <code> is — and the <code> was not a flex item, so upstream's flex: 1 1 auto; min-height: 0; overflow: auto on it never applied. The <pre> therefore kept the default min-height: auto ("never shorter than my content"), grew through the cap, and — the shell having no overflow of its own — painted the tail of the code over the prose below it. UAT 2026-10-03 17:00 measured it: shell 285px, <pre> 375px, <pre> overflow-y: visible, and the last rows of a Python block overlapped the "总结" paragraph. Fix: the override sheet makes the <pre> a flex column with min-height: 0. The <code> becomes a real flex item, upstream's scroll rule takes effect, and the overflow scrolls inside the block; the cap is untouched, and a block that fits lays out exactly as before. Rejected alternatives: moving overflow to the <pre> (every scrollbar rule, upstream's and this sheet's, keys on .codeblock-code, so the bars would be restyled or lost) and raising the cap (hides the defect, keeps long blocks unscrollable) |
webapp/styles/markdown-overrides.css, webapp/test/codeblock-overflow-containment.test.ts |
The overrides live in webapp/styles/markdown-overrides.css, a
webui-owned sheet loaded after styles/official-utilities.css
(app/layout.tsx); the vendored upstream sheet itself stays
byte-identical, because the desktop build shares it. Same-selector
rules there win by source order, which is why the import order is
load-bearing.
Assistant messages and Markdown file previews also render math formulas, in the same pipeline as Mermaid diagrams. Three input shapes are math; every other use of the dollar sign stays prose:
| Shape | Written as | Rendered as | Backed by |
|---|---|---|---|
| Inline | $E=mc^2$ |
KaTeX markup inside the paragraph | webapp/lib/math-renderer.ts (the marked webuiMath inline extension) |
| Display | $$\frac{a}{b}$$ |
A centred block (.katex-display) |
the same tokenizer, displayMode: true |
| Fence | ```math |
A centred block, dispatched through the language→renderer registry — the same seam mermaid uses, so neither fence language can shadow the other |
webapp/lib/math-renderer.ts (registerMathRenderer) |
| Aspect | Contract | Backed by |
|---|---|---|
| False positives | A single $ is math only when a closing $ exists, the body stays on one line, and the body does not start with a digit: costs $5 and $10, $HOME, an unclosed $\frac{ all stay prose |
webapp/lib/math-renderer.ts (tokenizer guard) |
| Invalid formula | An input KaTeX cannot parse degrades to the original source as code — inline/display shapes become <code class="inline-code">raw</code>, a ```math fence falls back to the plain codeblock shell (the registry's existing throw path). The rest of the document is unaffected; the page never blanks |
webapp/lib/math-renderer.ts, webapp/lib/markdown.ts (safeLanguageRenderer) |
| Sanitiser surface | KaTeX runs with output: "html" and emits only span, svg, path. The allowlist admits exactly those tags; svg keeps a fixed attribute set (xmlns, width, height, viewBox, preserveAspectRatio, class) with no href-like attribute, and <math>/MathML stays a DROP tag — which is precisely why HTML-only output is configured. Inline style survives only on span and only when the value clears isSafeStyleValue: no parentheses rules out url(...)/expression(...), and position/background/behavior are refused outright |
webapp/lib/markdown.ts (ALLOWED_TAGS, ALLOWED_ATTRS, isSafeStyleValue) |
| React tree | The style attribute reaches React as a parsed object (parseInlineStyle), because React rejects a string style prop outright — passing it through would silently drop all KaTeX layout |
webapp/lib/markdown.ts (parseInlineStyle), components/markdown-html.tsx |
| Trust | KaTeX trust stays false: \href renders as a red warning text node, never a link, so no URL can enter the DOM through a formula |
webapp/lib/math-renderer.ts (KATEX_OPTIONS) |
| Theme | Formulas are inheriting text plus CSS transforms; they need no per-theme re-render (unlike Mermaid, which repaints on the theme flip) and pick up both themes' text colours from the design tokens | webapp/styles/katex.css |
| CSS + fonts | webapp/styles/katex.css is vendored from katex/dist/katex.min.css (the same version as the katex devDependency) with the @font-face sources repointed at the vendored fonts in webapp/public/fonts/katex/ (60 font files + the MIT license notice). The stylesheet is loaded unconditionally from app/layout.tsx (~24 KB); fonts are served from /fonts/katex/… |
app/layout.tsx, webapp/styles/katex.css, webapp/public/fonts/katex/ |
| Bundle cost | katex JS is in the client bundle, not lazy-loaded the way Mermaid is: the math pipeline is synchronous string rendering (renderToString), and an inline $…$ can appear mid-sentence. Accepted as a known cost; the lazy-load lever exists only if bundle budgets demand it |
webapp/lib/math-renderer.ts |
Upgrading katex regenerates both halves in the same commit: replace the
files in webapp/public/fonts/katex/ from the new dist/fonts/, and
regenerate the stylesheet with
`sed 's|url(fonts/|url(/fonts/katex/|g' node_modules/katex/dist/katex.min.css
webapp/styles/katex.css`.
The katex dependency (0.18.7, MIT) is recorded in
release/dependency-licenses.json.
The preview component's header carries three controls, and the Markdown preview grows an outline panel. This slice also opens the webapp's first write path, so its boundary decisions are contracts, not defaults:
| Control | What it does | What it never does |
|---|---|---|
| ↻ refresh | Re-reads the file from disk, re-renders, and restores the scroll position. An external edit shows up on the next click | Never blanks on failure: a deleted/renamed file keeps the last content and shows a banner naming the likely cause |
| 预览/编辑 toggle | Flips a text preview (markdown, code) into a plain editor seeded with the disk bytes. Image and unsupported previews offer no editor | Never offered for credential-shaped paths without an explicit confirmation card (see below) |
| ✓ save | Writes the editor buffer through POST /api/fs/write, then shows "saved at HH:MM" and re-renders the preview from the saved bytes |
Never automatic, never silent-on-failure: the buffer survives a failed save and the server's reason is shown |
The endpoint is the only write surface the preview opens, and each invariant is enforced server-side — the webapp is a presenter over the structured answer:
- Containment — the same
assertWorkspacePathgate every other/api/fs/*route funnels through (realpath-resolved, symlink-aware). The write path adds no escape hatch. - Credential guard (slice 16 alignment) —
.env/*.pem/id_rsa/credentials*… default-refuse with403 {code:"credential", credentialReason}.confirm:truereleases the write and emits the samecredential.overridestderr audit line as the read override (endpoint:"write"). Reason: the server broadcasts a LAN URL, and a web-editable.envmakes every LAN peer an author of the local machine's config. - Conflict detection — the read (
GET /api/fs/read-file) carries the file'smtime; the save sends the(expectedMtime, expectedSize)pair it recorded when the file was opened. Either value drifting from the live stat answers409 {code:"conflict", diskMtime, diskSize}and nothing is written. The panel shows a conflict card with two explicit exits: overwrite the disk version, or load the disk version. A save that omits the baseline is the explicit-overwrite shape the panel only sends after the user answered that card. - Controlled write — the handler is a bare
writeFileSync(path, content)on the gated path. No shell, no exec, no command interpolation anywhere on the path. - Caps — bodies over 512 KiB (the same figure the read caps at)
answer
413 {code:"too-large"}; the editor edits existing files only (a vanished file is surfaced, never created).
The outline is derived from the rendered DOM — never a second parse
of the markdown source — so the outline lists, by construction, what the
page shows. Heading ids are assigned onto those DOM nodes (stable slugs,
-2/-3 suffixes on duplicates); a click preventDefaults the anchor
and smooth-scrolls the heading into view via scrollIntoView, so the
jump animates and the scroll position still persists through the tab's
usual channel. The panel is sticky within the scroll viewport (its
max-height pinned to the viewport's height, so a long outline cannot
outgrow the pane it floats in) and tracks the scroll position to
highlight the current section. Documents without headings render no
panel (no empty box), Mermaid diagrams never enter the outline (a
diagram is not a chapter), and entries are readable in both themes
through the design tokens. Below ~300px of content width the outline
hides rather than squeezing the document — the preview column's own
minimum (320px) still shows it.
The "process" half of an assistant turn is carried by two native
<details> disclosures. The components live in
components/activity-group.tsx (lifted out of chat.tsx for the same
reason U8 lifted loading-states: the SSR render tests can load the
module without chat.tsx's @/-aliased import graph). All data comes
from the existing transcript decode (groupActivity in
webapp/lib/transcript.ts); the server transport is untouched by this
ticket.
Activity group (one run of adjacent thinking/tool steps; the summary reads like "Thought 1 time, ran 1 command"):
- The summary row IS a
<summary>: one click anywhere on the row toggles the group (keyboard reachable), replacing the previous text-button + separate-chevron-button pair. - The
file-editcontribution counts distinct FILES, not edit calls. Every other category counts the calls in its bucket, because every other sentence is a call sentence (「执行 1 条命令」). 「已编辑 N 个文件」 has a file as its subject, so a model that edited one file five times must not read as five files — that number used to disagree with the turn's edited-files card, which renders the same sentence fromcollectEditedFiles. Both now de-duplicate through one key,editedFileKeyinwebapp/lib/transcript.ts(separators folded, case left alone), so a group header and the card cannot report different N for one turn. An edit call that named no path contributes nofile-editline at all rather than a phantom file. How many calls ran is not lost: it is the summary'stools, which the turn bar reports as 「用了 N 次工具」. - The expanded body carries a 1px timeline spine on its left edge
(
.timeline-spine,border_light, ported parameters). - While the run holds a tool whose status has not settled (no
[completed]/[failed]line yet, or an explicit[in_progress]), the group carriesdata-active="true"and cannot be collapsed — a click during the run snaps straight back open; only a settled turn folds. The predicateisActivityGroupActiveis an exported pure function. - Reachability of
data-active(stated plainly): protocol-level probes (live engine turns, 120ms sampling, including asleep 15tool and a 15-second streaming-output tool) show that under the current ACP transport the tool's→ nameheader and its[completed]status line land in the SAME frame — the engine emits thetool_callnotification (carryingupdate) only at completion, with no incrementaltool_updatein between, and the server'sapplyToolUpdate(mcode-acp.js) defaults an update without a status field to[completed]. The "running tool block" intermediate state therefore does NOT occur under the current engine transport:data-activeforced-open is a capability reserved for the engine emitting tool-start events / non-terminal intermediate updates in the future. The decode contract (no status line = running), the front-end predicate and the snap-back logic are all in place and pinned by unit tests, so the behaviour activates the moment the engine sends the events. The SSE snapshot coalescing window (default 16ms) is not the masking cause. - Default-open follows the upstream orchestration (
AssistantBody'sexpandProcessByDefault+renderActivityParts): a mixed run (thoughts AND tools) opens expanded; a pure-tool run starts collapsed; thinking rows nested in a mixed run start collapsed while a thoughts-only run starts with its thinking row expanded. - The group body is uncapped and has no scrollbar of its own (webui-parity
61, restoring the desktop shape): the body used to carry
max-h-[230px] overflow-y-auto, so a long tool run nested a second scrollbar inside a page that already scrolls and cut its own steps off mid-list. The reference stylesheet gives.activity-group-itemsgap: 0plus a 28px minimum row height and nomax-height; this repository now matches — zero gap as the component'sgap-0, and the 28px floor inapp/globals.cssas.activity-group-items > *(outside every@layer, so Tailwind's purge cannot reach a hand-written rule). Removing the cap costs no render work either way: a collapsed<details>keeps its body in the DOM, so the row count was never the reason for the cap. The one scroll container the desktop keeps inside a group is the tool detail'spre, capped at 180px (the reference.webui-tool-detail-section pre); this repository shipped 320px.
Thinking block (one thought):
- The summary row reads icon + status copy + elapsed seconds + chevron. While streaming it shows 「推理中...」 with the second counter ticking every second; when the turn settles it becomes 「已完成推理」+ the frozen total.
- The expanded body renders through the existing Markdown pipeline
(
lib/markdown.ts; the KaTeX and mermaid language renderers register with the component), not as plain text. - A body taller than 224px is clamped (
.is-clampedplus a bottom gradient mask) behind an 「展开 / 收起」 toggle. - Streaming keeps the block expanded; the end of the turn collapses it again (unless the user had opened it by hand).
Duration data boundary (stated plainly): the line-oriented transcript
carries no per-thought timestamps, so the seconds tick from the snapshot's
running.startedAt (the turn's start) — the same turn-level anchor
upstream feeds WebuiThinkingBlock as processingStartedAtMs. The
streaming seconds therefore read as "time elapsed in this turn", and the
frozen total is the turn-elapsed value at the moment that thought settled
— NOT an engine-measured per-thought duration (the engine does not expose
one). A cold-loaded historical thought has no anchor and omits the seconds
rather than inventing them.
Showing the seconds additionally requires the thought's streaming window ("the tail block is a thought") to be exposed in at least one SSE snapshot frame. Two boundaries confirmed on the live instance:
- Swallowed streaming window: a very short thought, or one that lands in the same frame as the turn's end, may never appear as the tail block in any snapshot frame; that thought then freezes with no seconds (「已完成推理」 with no number). The data was already merged at the transport layer — the renderer cannot reconstruct it afterwards.
- Row remounting (the fixed primary cause): mid-turn, tool headers
land only at completion and prose lines stream in between, so the
activity runs are re-cut frame by frame; the thinking rows originally
keyed by their within-group position remounted on every re-cut, wiping
the elapsed state exactly as the turn settled (reproduced live: ticking
1s→4s, zeroed at finalize). Fix:
assignActivityBlockKeys(an exported pure function inactivity-group.tsx) assigns global birth-order keys, and the same scenario now keeps its seconds. Rare line reorderings (e.g. a block displaced by a late-arriving prose line) can still drop an individual thought's seconds.
The streaming verdict itself is derived, not signalled: the tail unit of
the transcript is an activity run whose last block is a thought
(streamingActivityIndex in chat.tsx). The transcript has no
thinking-level streaming marker (the ▍ cursor only marks the trailing
assistant block), so this is the most honest signal the render layer can
derive.
Invariant: while a turn streams, the tail of the transcript is inside the
scroller's visible box. Every token the engine emits must be readable without
the reader scrolling. A turn appends a block and then grows it token by token,
so the DOM below the reader grows while their scrollTop does not move; with
nothing pinning the offset, the answer is laid out below the fold and the reader
watches the thinking indicator for the whole turn. Reproduced on the unfixed
build: over a 15-second turn scrollTop never left 0 while scrollHeight went
688 → 1196.
This is not virtualization. Below VIRTUAL_LIST_THRESHOLD (200 units) the
window is useVirtual: false and every unit is in the DOM; the measured turn
peaked at 8 units. The tail was rendered and simply off-screen.
| Concern | Decision | Rejected alternative |
|---|---|---|
| Who owns the offset | Exactly three: the tail follow, the reader, the persisted-position restore. Browser scroll anchoring is switched OFF on the scroller (.chat-scroll, app/globals.css) because it is a fourth, implicit owner — it adjusted scrollTop by 55 → 125 px on its own while the transcript grew above the viewport, which the follow reads as the reader leaving |
Leaving anchoring on. A transcript that grows at the tail needs pinning, not anchoring; with it on the follow cannot attribute a move to its cause |
| When the follow turns off | When the container moves off the pin the follow itself last wrote (isAwayFromPin, 2 px tolerance). Content is appended BELOW the reader, so growth alone never changes scrollTop |
A nearness test (isNearBottom) at scroll-event time. Scroll events dispatch asynchronously, so by the time the handler runs the next SSE frame may already have grown the transcript and the test answers "no" for a reader who never left. Shipped once, measured latching off ~20 s into every turn, replaced |
| When the follow re-arms | The reader scrolls back to the tail, or clicks the "jump to latest" pill (which calls followNow() — the pill's own scroll is about to change the position a re-arm would read) |
Re-deriving from metrics every commit, which cannot tell "the reader moved" from "the transcript grew" |
| Reader scrolled up | Left exactly where they are, for the rest of the turn; the pill stays available. Measured: 0 px drift across 157 samples / 40 s of continued streaming | Following anyway, on the argument that the reader will want the answer eventually. Dragging a reader who is reading history is the other half of the same annoyance |
| While a turn is live | The persisted-position restore is gated off (sessionRunning). initialScrollTop is re-read from storage on every render, so the restore re-armed itself from the follow's own persisted writes and dragged the container back one frame after each pin |
Letting both write. Two owners for one offset is the defect, not the fix |
| Animation | Instant, not smooth — a smooth scroll chases a target that moves with every token, so it lags the stream and overshoots at turn end | behavior: "smooth" on the pin |
| Virtualization | Untouched. The follow is orthogonal: one number written per commit, before paint, in a layout effect | Disabling virtualization above 200 units to keep the tail rendered. The tail was always rendered |
The restore is unaffected for a settled session — that is the case it exists for. Reopening a session left on the tail lands on the tail; one left mid-history lands mid-history.
One tool call renders as a native <details> card (ToolCard), on the
structure of the reference WebuiToolRow:
- Human labels: the summary row shows the human tool name
(
bash→ 终端,read→ 读取文件,edit→ 编辑文件,grep→ 搜索, …) fromtoolCallLabelinwebapp/lib/tool-projection.ts— a bilingual port of the reference desktop copy (the zh column is the reference copy verbatim, the en column its English counterpart; both locales ship at equal weight). Unknown names fall back to 「工具 / Tool」, the reference fallback. Reference entries carrying{{name}}template placeholders (read_skill_file& co.) are not carried over: this renderer has no substitution site, so keeping them would display the raw placeholder. - Five-state normalisation:
normalizeToolStatusmaps the wire status line onto 等待中 / 运行中 / 已完成 / 失败 / 已取消 (plusunknown), with the reference desktop copy; Desktop's numeric codes 1/2/3/4/5 map to running/completed/error/pending/pending and are kept for parity (today's wire only writes string statuses). One deliberate delta: a missing status line counts as running (the reference records unknown), because the wire writes the→ nameheader first and the status line only when the call settles — "no status" means in-flight on this transport, the same ruleisActivityGroupActiveencodes. A completed call renders no status chip (reference rule — settled rows stay clean). - Three body sections: 输入 / 结果 / 错误. The input args moved OFF
the summary row into the 输入 section (orchestration decision); a
failed call's output lines render as the red 错误 section (the wire
writes failure text as ordinary output lines — there is no separate
error field), falling back to 「执行失败」 when a failure has no
output; any section longer than 2000 characters clamps with
...(clampDetailText). A running call with no output yet shows 「运行中…」. - Read-style resource paths: for
read/read_filethe resource path is lifted onto the summary row (basename displayed, full path on thetitleattribute,toolSummaryResourcePath). The args derivation comes first (the reference rule), falling back to the@ pathlines the decoder collects intotoolPaths. Live sessions show the engine writes read calls with JSON args on the header (→ read {"path": …}, keypath) — real traffic takes the first-priority branch and the summary path renders (verified 3/3 read cards on a live instance, each title the full absolute path). A header-without-args spelling also occurs; there, when the@ pathlocation lines are separated from the tool body by a blank line and the 「N more lines」 truncation marker, the decoder's existing orphan rule drops them (which also correctly discards the next tool's status lines mixed in among them — the rule is right and stays) and the summary path is absent — a data-source shape difference, not a rendering defect. No other tool lifts a path. - Icons: a 16×16 SVG catalog (
components/tool-icon.tsx) replaces the unicode-glyph placeholder. Paths are transcribed from the referenceWebuiToolIconregistry; this wire's icon vocabulary (SummaryIconType) is not name-identical with the reference categories, so a few entries use the nearest reference glyph (plugin→ reference combine,file-edit/edit→ reference code,agent→ reference bot,skill→ reference task), and the two types with no reference counterpart (summary,alert) are drawn here in the same style (viewBox 16, stroke 1.25, round joins/caps). On failure the icon and the status chip turn error-coloured. - Legality fix (QA-registered): the old card header was a
<button>with the subagent badge — another<button>— nested inside it: invalid HTML. The<details>/<summary>rebuild removes the nesting (a<summary>may legally contain interactive descendants). The orphandata-message-collapse-triggerattribute (no consumer anywhere in the repo) is deleted.
No per-tool duration, by decision: the reference tool row summary reads "name · status" only — durations live on the turn bar and the thinking block. That is the orchestration call, not an omission.
While a turn is in flight, a process bar renders at the transcript tail
(TurnProcessDisclosure, lifted from chat.tsx into
activity-group.tsx and rebuilt on the reference WebuiTurnProcess):
- The summary row is the composite 「思考 N 次,用了 M 次工具,已执行
N 秒」 (「…共执行 X 分 Y 秒」 once the turn settles); zero-count parts
drop out; durations over a minute read 「X 分 Y 秒」, under it bare
「N 秒」. The counting matches the activity group (adjacent thinking
blocks merge), and the span is everything between the previous user
message and this turn's end (
webapp/lib/turn-stats.ts; the forward and backward scans share the one rule). - The settled state shows the output rate
N token/son the right. The figure is an estimate: the wire transcript carries no per-turn token count (the ACPusageevent only accumulates session totals server-side, and this ticket's red line forbids touching the four server files), so it uses the same fallback the reference applies when its runtime reports nousage.outputTokens:answer characters / seconds. Character-to-token ratios differ between scripts — read it as an order of magnitude, not a meter. - The tick is effect-driven, so the SSR and hydration first frame deterministically render 0 seconds (no server/client markup divergence).
- The settled bar is transient (stated plainly): when the turn
ends, the bar briefly flips to 「共执行 X 分 Y 秒」 and gains the
N token/sfigure, but that settled state is visible only for about one SSE snapshot window (measured on the order of 150ms); after the session stream finishes, and after any reload, it is gone. Root cause is the existing server link: finalize pushes the§§ processed_durationmarker into the in-memorycs.chat(the SSE briefly delivers it), but the marker does not survive the session state persistence/rebuild, so the front-end decode has noprocessedDurationto render. This ticket's red line forbids the four server files; fixing it (persisting the marker, or moving it to a structured field) needs its own ticket — until then the settled bar is transient by design of the transport. - A 0.5px separator closes the bar from below; the
turn-process-disclosuretestid is kept.
A settled turn bar carries the desktop's > chevron
(turn-process-chevron). Ticket 46 removed it on the grounds that there
was "nothing to expand" — this repository lays the thinking and tool
steps out flat as activity groups instead of nesting them inside the
turn bar's collapsible region, so the chevron had nothing to point at.
Webui-parity 61 changes that premise: the chevron is state
coordination that drives the turn's activity groups, without
re-arranging the DOM.
- When it appears: the turn has thoughts or tool calls. The predicate
is the exported pure function
hasExpandableTurnContent(stats)— the same two counts the composite summary already prints, and the same question the reference asks throughWebuiTurnProcess'shasExpandableContent(whichAssistantBody.tsxcomputes from that turn's thinking text, tools and activity segments). A plain question turn (no thought, no tool) has nothing to disclose and renders the bare summary line with no chevron; that is the desktop's own rule, and the reason the turn in desktop screenshot 02 does show one is that the turn itself had process content. - A live turn never shows it. The reference reaches the same state
through its
forceExpanded/disabledprops, whose docblock says those modes "suppress the toggle and keep available details open". Here there is a harder reason as well: an activity group holding a running tool is force-open underdata-activeand cannot be collapsed, so a live chevron would be a control that cannot act. - What expanding shows: the turn's thinking and tool steps — the
activity groups. The turn's answer text is not part of it: the
reference keeps the answer outside the collapse through
collapsedContent, and here the answer is never inside a group. - The coordination key is the turn ordinal, not the unit index.
Activity runs are re-cut on every streaming frame, so a unit index
drifts while its turn ordinal only advances when a new user prompt
arrives; keying on units would detach the chevron from the groups it
drives the moment a tool header lands.
computeTurnLayoutinwebapp/lib/turn-stats.tsis the pure function that produces the ordinals, and the intent lives in oneMap<turn ordinal, expanded>inChat. - Until the user clicks, every group keeps its own default (mixed run
open, pure-tool run collapsed). The chevron's
aria-expandedreadscomputeTurnLayout#defaultExpandedByTurn(true when any group of the turn opens by default) and inverts it on the first click, so the first click always changes something visible. After that the turn's groups move as one block. - Neither the
data-activeforced-open rule nor a group's own folding semantics change: a collapse intent never overrides forced open.
The last block of a turn is the edited-files card — the one in the
desktop screenshot 06-browser-tree-tasks-review.jpg. Ticket 77 mounted
it once, at the end of the transcript, from the file paths the turn's
edit tools named, and recorded in this document that the desktop's line
counts and undo button were unreachable. Both became reachable in
ticket 83, and this section now describes what the card does with
them.
The header reuses the existing activity.editedFiles key for its
sentence; ticket 83 added keys only for the new affordances (undo, redo,
the failure copy). The activity-group summary and this card are the same
sentence, so they have to stay the same sentence — and, since they are the
same sentence, the same number: the group header counts distinct files
through the same editedFileKey (see the activity-group section above).
Two sources, in strict order of authority, and never merged:
| Source | What it is | What it can prove |
|---|---|---|
| The engine's record | GET /api/turn-diff with the turn's assistantMessageId |
The real per-file +N / -N, the real file list (which includes edits made through tools whose arguments name no path), and the engine's own canUndo / canReapply |
| The transcript scan | collectEditedFilesByTurn (webapp/lib/edited-files.ts), keyed by the layout turn ordinal |
Only the file paths a file-edit tool named. No counts, no gates. |
The engine's list REPLACES the scan rather than merging with it. A merge
would double-count a file the two name differently, and the counts would
then sit on the wrong row. When there is no record — a session recorded
before the turn coordinate shipped, a legacy transcript read, the exec
transport — the scan is the whole story and the card is exactly what
ticket 77 shipped.
The engine persists a turn's record under the msg_id of that turn's
last assistant message (local-runtime-v2/.../turn-outcome.ts reads
the last agent message response; a turn carries more than one id on
the wire, one per message segment). The ACP transport already delivered
that id; the server used to drop it.
- Live:
acp.mjs#promptkeeps the lastagent_message_chunkmessageId;mcode-acp.js#finalizewrites§§ turn_msg=<id>next to the§§ processed_duration=Nmsmarker it already wrote. The§§family is the established convention for server-written per-turn metadata — third reuse, not a new grammar. - Restored:
server/lib/transcript.js's v2 probe now selects theturn_idandmsg_idcolumns the message table already carries, and synthesises the same marker from the last assistant row of each turn. Switching away and back gives an existing session its coordinates without asking the engine anything.
decodeTranscript consumes the marker and hangs the id on the turn's
LAST assistant block. A transcript without the marker decodes exactly as
it did before — every session older than the marker has to keep
rendering, and a decoder that assumed the marker would throw on all of
them.
| Desktop element | Here | Why |
|---|---|---|
| Header glyph | Present | The pencil icon from the existing icons.tsx set. |
| 「已编辑 N 个文件」 | Present | N counts distinct files, not edit calls. With a record it is the engine's own file count; without one it is the count of paths the scan could name. |
Green +N / red -N |
Present, with a record | The engine's own per-turn counts, summed in the header. Absent without one — a card showing +0 -0 would say "this file did not change", which is a different and false claim. A zero side is not drawn as +0 / -0 either. |
| 「撤销」 (undo) | Present, gated on canUndo |
Calls POST /api/turn-diff/revert, which rewrites real workspace files. |
| 「重做」 (redo) | Present, gated on canReapply |
Calls POST /api/turn-diff/reapply. Appears independently of undo: a reverted turn offers redo and no undo. |
| 「Review」 | Absent — the file rows replace it | A file row opens the real preview through the existing onOpenFile chain (red line 4). |
| File row: type icon + name | Present | A file icon plus the path's trailing segment; the full path rides on title and data-file-path. |
| Per-row added/deleted | Present, with a record | The engine's per-file counts, not the header total repeated. |
| Collapse (3 rows, then expand) | Present | A pure client-side state machine, reduceEditedFilesCardState; the toggle appears only past three rows. |
The two buttons are not greyed out when they cannot act; they are
absent. Only the latest turn diff can be changed — the engine answers
canUndo:false before the user clicks, and 409 TURN_DIFF_CONFLICT
if they click anyway — so a disabled button would promise an action the
engine has already refused. The card never re-derives "is this the last
turn?" from the transcript; the whole coordinate system exists because
that inference cannot be made safely.
A revert rewrites files the browser is already showing, so five things move, in this order:
- the server's session-tree cache (
invalidateSessionTree()); - a
session-tree-changedbroadcast (the sidebar); - a
workspace-files-changedbroadcast — a new named SSE frame, no payload, the only signal the webui has that files on disk moved; - on that frame, the files tree re-reads every directory it has open and the git panel re-reads status and branches;
- on the same frame, the open file preview re-reads through the refresh path — scroll position preserved, a deleted file shown as a named banner rather than a blank pane, and a dirty draft left alone, because its baseline is what a later save conflict-checks against.
- The collapse state is not persisted in
localStorage: a reload returns to collapsed. It is derived state, not a user preference, so red line 3 is untouched. - The coordinates, the scan and the records are all derived from the
full
unitslist, never from the virtualisedvisibleUnitswindow, so a turn outside the window is still fetched and its numbers do not change under the user. The card itself, however, is a list child: it renders inside thevisibleUnitsloop at its own turn's last unit. AboveVIRTUAL_LIST_THRESHOLD(200 units) a turn's card therefore appears and disappears with the scroll window, the same as that turn's messages do. This is the one place the card is windowed, and it is a deliberate trade — a card pinned outside the window would render in a position the reader cannot see. Non-virtualised transcripts (≤ 200 units, the overwhelming majority) render every turn's card unconditionally. - Each turn's card sits at its own turn's last unit, so a three-turn session reads as three cards. Ticket 77's single session-wide card at the transcript tail is gone: the trailing message-action row is now the last element in the column.
- With no
onOpenFilewired, a file row degrades to plain text rather than to a button that does nothing. - The route exposes
applications.session.diffand nothing else; theapplicationstree also carriessession.lifecycle, which can delete a session. previewStateis declared by the protocol and is alwaysundefinedon this path. Nothing renders it.
The two waiting windows on the conversation surface have distinct treatments, and neither is a bare spinner:
| Waiting for | What renders | Where it lives |
|---|---|---|
| The first state snapshot (page load, engine boot) | TranscriptSkeleton — a shimmer placeholder shaped like the real transcript: right-aligned user bubbles, full-width assistant text lines, a tool-summary row with indented output lines. The connection copy (connecting / disconnected) stays underneath |
The page-level !state branch in app/page.tsx; the component itself is webapp/components/loading-states.tsx |
The current turn's output (running.active) |
ActivityPulse at the transcript tail — the desktop's three-dot loader plus a shimmer bar sitting where the next line of output will land. The label shows the phase the engine reported (thinking / working / …) first, then rotates through the desktop's weighted phrase table |
components/chat.tsx#ThinkingIndicator, gated by the exported pure function isSessionActivityActive |
Invariants worth keeping when touching either branch:
- The skeleton's bar colour is the user-bubble token (
--bg_grouped_tertiary) with the 15%-black overlay token as the sweep, so both themes come from the token layer and the placeholder sits in the same colour family as the rows it stands in for. No per-theme rules exist. prefers-reduced-motion: reduceswitches every animated class off explicitly (.mavis-skeleton-bar,.mavis-loading .mavis-dot) inapp/globals.css, on top of the generic duration catch-all. The dots' rule carries!important:styles/official-utilities.cssloads afterglobals.cssand re-declaresanimation-nameon.mavis-dot-a/b/cat the same specificity, so a plain declaration would lose the cascade and the dots would only look still through the duration hack. The label text survives with the motion removed.- Session switching does not show the skeleton:
POST /api/sessions/switchresolves the full transcript before responding and the next SSE snapshot replaces the old content wholesale. The skeleton's only trigger is a missing snapshot, which is the cold-load path. - Rendering tests for both components and the reduced-motion tripwire live in
webapp/test/loading-skeleton.test.ts(SSR throughrenderToStaticMarkup; interaction-level coverage of this component can now use the DOM harness — see the tests section). - Streaming-label phrase rotation (webui-parity 61, restoring the desktop
shape): the desktop does not park one static label on screen for the
length of a turn. The schedule and the draw are transcribed from the
reference
ActivityIndicator.tsx: 2000–3000 ms (drawn) before the first swap, 3500 ms between swaps after that, a weighted bucket draw — basic 0.75, specific 0.15, motion 0.1 — uniform inside the drawn bucket, with the previous phrase filtered out so nothing repeats back to back. The tables are inwebapp/lib/thinking-phrases.ts, indexed byLocalewith the same structure in both languages; a separate module rather than the flatlib/i18n.tsdictionary, for the reasonlib/i18n-agent-team.tsgives. The draw and the schedule are the exported pure functionspickWeightedPhrase/computeThinkingPhraseStartDelay, and the timer lives inActivityPulsealone. - Under
prefers-reduced-motionthe rotation continues, deliberately. A phrase swap is a discrete text replacement, not motion: there is no translation, no scaling and no continuous movement for a vestibular trigger to react to, and the desktop makes the same trade — its reduced-motion branch halts the lottie and leaves the label ticking. Stopping the rotation would re-install the frozen-label defect G5 exists to remove. The animated half of the indicator (three dots, shimmer) is still switched off by the explicitglobals.cssrules above, so a reduce-motion user sees a still indicator with a still-ticking label. - No side effects on streaming: the rotation's state lives in
ActivityPulseitself, so a swap re-renders that one<span>(activity-indicator-label) and nothing above it — the transcript bodies, the markdown and the streaming cursor are not in the update path. The effect depends on the phrase table alone (a module-level constant looked up by locale, stable identity), never on a per-render closure and never onlabel— the phase copy changes with the engine, and depending on it would reset the schedule mid-turn. The timer is a chainedsetTimeoutrather than asetIntervaland is cleared on cleanup; the component unmounts when the turn settles, leaving no pending timer. Effects do not run on the server, so the SSR and hydration first frames are both the phase copy — no first-paint swap. - The streaming cursor blinks; it does not breathe (webui-parity 61, G6).
The trailing cursor glyph on a streaming assistant block is drawn with
.stream-cursor, a square wave:stream-cursor-blink 1.1s steps(1, end) infiniteover two opacity levels, lit for 60% of the cycle and dim for the rest. It replaced Tailwind's stockanimate-pulse—pulse 2s cubic-bezier(0.4, 0, 0.6, 1) infiniteover@keyframes pulse { 50% { opacity: .5 } }— a symmetric eased breath with no instant edge and no off state: a loading tell, not a caret. Four choices are load-bearing.steps(1, end)holds each level for its whole interval and jumps at the interval end, which is what makes it read as a blink. The 1.1s period is faster than the 3.5s phrase rotation and not a ratio of it (3.5 / 1.1 = 3.18), so the two rhythms never lock into one slow compound beat. The dark level is 0.2 rather than 0, because this glyph sits at the end of a prose run and owns no cell of its own the way a terminal block cursor does — at 0 the paragraph looks like it lost its last character for half a second. The60.01%stop exists so the dark level has a declared start and the edge is a jump rather than a fade across 60% → 100%. The rhythm is deliberately not keyed to token arrival. The cursor span is reconciled as the same element for the whole turn, so the animation free-runs; a "flash on each push" behaviour would need a per-token remount, which would restart the animation before it ever reached its dark phase — a solid glyph exactly while the rhythm matters most — and churn one DOM node per token for nothing. Underprefers-reduced-motionthe class is switched off by an explicit rule in the sharedglobals.cssblock, like every other animated class there, and the caret renders solid and still: the glyph is the only carrier of "where the next token lands", so a cursor that vanished under reduce-motion would break the insertion point. Pinned bywebapp/test/stream-cursor.test.ts. - Boundary with ticket 46: while a thought streams, the 「推理中...」+
ticking-seconds readout lives on the thinking block's summary row inside
the tail activity group, and the live 「已执行 N 秒」 turn bar sits at
the transcript tail (see the "Session rendering" section above); the
ActivityPulsein this table (three dots + shimmer + phase label + rotating phrases) also appears only at the transcript tail, below the turn bar. Different positions, different jobs; none replaces another.
| Key | Channel | Owner | Introduced by | Shape |
|---|---|---|---|---|
webui:ui:v1:<cid> |
localStorage |
webapp/lib/persist.ts#uiStateKey |
slice 07 (reopen state) | {version:1, cid, state:{panel, panelTab, sidebarCollapsed, lastSessionId, appearance}} — appearance (slice 18) is the three-state picker choice ("light" | "dark" | "system"); applyAppearance writes through this envelope |
webui:scroll:v1:<cid>:<sessionId> |
localStorage |
webapp/lib/persist.ts#scrollKey |
slice 07 | {version:1, cid, sessionId, scrollTop, savedAt} |
webui:workspace-tabs:v1:<cid> |
localStorage |
webapp/lib/persist.ts#workspaceTabsKey |
slice 15 (workspace columns) | version-discriminated state (WORKSPACE_TABS_VERSION) — see lib/workspace-tabs-state.ts |
webui:open-file:path |
localStorage |
webapp/lib/open-file.ts#STORAGE_KEY |
slice 12 (file preview) | bare path string or absent |
webui:files-tree:<workspaceDir> |
sessionStorage |
webapp/components/panels.tsx (slice 01) |
slice 01 (file tree) | {version:1, workspace, expanded[], filter, showHidden} |
file_open_in_new_tab |
localStorage |
webapp/lib/settings-local.ts |
ticket 48 (settings General page) | bare "true"|"false" string; deliberately outside the webui: namespace — same key and format as the desktop reference so one browser profile shares the preference across both clients. Default "true" here (reference: "false"); read by app/page.tsx#openFileTab |
file_line_wrap |
localStorage |
webapp/lib/settings-local.ts |
tickets 48 + 52 | bare "true"|"false" string, reference-shared namespace; default "true"; read per mount by components/code-view.tsx (code-file previews) and components/markdown-html.tsx (markdown codeblocks: chat, activity groups, file previews) |
webui-context-window-usage |
localStorage |
webapp/lib/settings-local.ts |
ticket 48 | bare "true"|"false" string, reference-shared namespace; default "false"; read at mount and followed live by components/context-meter.tsx through subscribeContextWindowUsage |
webui-follow-up-behavior |
localStorage |
webapp/lib/settings-local.ts |
ticket 48 / SB-4 | bare "off"|"queue"|"steer" string (anything else reads as "queue"), reference-shared namespace; read by components/composer.tsx and republished on every write |
webui-desktop-notifications |
localStorage |
webapp/lib/settings-local.ts |
SB-9 | bare "true"|"false" string; default "false"; written by the General page's 桌面通知 row, published on every write for subscribeDesktopNotifications. Stores the user's intent only — the browser's Notification.permission is read live from webapp/lib/desktop-notify.ts and is never persisted |
webui-shortcut-bindings |
localStorage |
webapp/lib/shortcuts.ts |
ticket 55c (settings Shortcuts page) | {"global-search":"Ctrl+Shift+P", …} — rebindings of the live shortcut rows only, written when the user records a new combination and removed entirely when the last one is cleared. Re-validated against the registry on read: a stored id that is no longer dispatched, or a chord that no longer parses, is dropped rather than honoured, so a hand-edited entry cannot widen what the page dispatches. Read at every keydown by app/page.tsx (through effectiveBindings) and once per mount by the settings page |
webui:project-custom:v1 |
localStorage |
webapp/lib/project-custom.ts |
ticket 55c (project context menu) | {version:1, titles:{<projectKey>:<customName>}, pinned:[<projectKey>]}. Deliberately not cid-namespaced: a rename or a pin describes the project, not a browser session, so every tab of this browser shares it. Best-effort write, silent failure; a project's entries are cleared when its remove completed with every session deleted |
Except for ticket 48's four reference-shared keys (file_open_in_new_tab,
file_line_wrap, webui-context-window-usage, webui-follow-up-behavior,
which deliberately use the desktop reference's bare key names),
all keys share the webui: prefix and are best-effort writes (debounced
150 ms for ui and workspace-tabs; immediate for the others). A failed
write leaves the in-memory state correct and the persistence silent — the
failure mode we care about is the app/global-error.tsx crash, not a quota
error here. Per-session scroll keys are deliberate: a refresh restores
the user's place in each conversation independently.
One timing invariant guards all of it (webui-parity 106): the page root
never reads these keys during render. The prerendered server HTML and the
client's first (hydration) render must be identical, and a render-phase
storage read breaks that equality the moment the state === null skeleton
changes shape. app/page.tsx renders its first frame from the shared
DEFAULT constants and applies the stored payload in one post-mount effect;
the three write-back mirrors are gated on that restore having run, so the
defaults-seeded first render cannot overwrite the stored payload. What the
user sees is unchanged: the skeleton is still up while the restore lands,
and by the time the first snapshot arrives the saved layout is in place.
A /-prefixed line in the composer is not automatically a command. Two
endpoints can consume one, they implement different sets, and the
composer decides between them before anything is sent.
| Input | Endpoint | Why |
|---|---|---|
/new /clear /status /sessions /review /help /usage /stop — bare, no argument |
POST /api/cmd |
the button-command set; /api/cmd claims exactly these eight |
/goal <text>, /goal-done, /goal-blocked |
POST /api/send |
typed webui commands, implemented by handleLocalSlash; /goal needs its argument and has no /api/cmd equivalent |
/compact and every other engine command |
POST /api/send |
handleLocalSlash's default branch forwards the line to mcode unchanged, which is how engine commands work at all |
| anything unclaimed | POST /api/send |
same forward; the engine answers in the transcript |
/clear now (a claimed command with an argument) |
POST /api/send |
handleCmdCommand matches the whole text after the slash, so the argument makes it a different string; handleLocalSlash parses the name and runs the same gated handler |
The /api/send side is not a disjoint list.
server/lib/interaction/command-registry.js declares
SEND_SLASH_COMMANDS (goal, goal-done, goal-blocked, clear,
new, status, review, help, usage) — nine names, six of which
(clear, new, status, review, help, usage) are also button
commands. handleLocalSlash consumes them, which is why the
/api/cmd 400 branch asks isSendSlashCommand(name) first and
phrases its suggestion as "send it as a normal message" for those.
Routing still prefers /api/cmd for the bare form;
SEND_SLASH_COMMANDS never influences routing.
The decision is routeSlashInput (webapp/lib/slash-routing.ts), called
from composer.tsx#submit. The set it compares against is declared once
on the server in server/lib/interaction/command-registry.js
(CMD_BUTTON_COMMANDS, read by the /api/cmd 400 branch and by
/help's fallback); the browser carries a mirror because the bundle
cannot import a server module, and webapp/test/slash-routing.test.ts
compares the mirror against the registry and against the case labels
parsed out of both dispatchers in interaction/commands.js. A command
added on one side alone fails the gate.
While the composer holds a single /-word and at least one command
matches, the palette is open. The keys mean this, and the meaning does
not depend on how many candidates are showing:
| Key | Effect |
|---|---|
Enter |
sends what is in the box, palette open or not |
Tab |
accepts the highlighted candidate into the box; the box is not sent |
↑ / ↓ |
moves the highlight (wraps) |
| clicking a row | accepts that row |
Esc |
clears the draft, including a palette that was open |
The candidate count is deliberately not an input to any of this.
availableCommands reports every command in two groups — the engine's
own mcode list and the webui button list — and a name can therefore
arrive more than once, so a fully typed /status reaches the composer
as two identical candidates. A rule of the form "Enter completes while
the list is ambiguous" therefore fired on an unambiguous command and
swallowed the keystroke that was supposed to run it: the box kept its
text, the command never ran, and the next Enter sent the bare word as
a chat message. The decision lives in shouldCompleteSlashWord
(webapp/lib/slash-routing.ts), which takes the key and nothing else.
Dedupe is applied on the way to the screen, not inside that decision —
see below — because the two defects are independent and neither is the
other's precondition.
A second defect shared that key handler and is fixed with it:
availableCommands carries bare names (name: "status"), so
writing a candidate back verbatim produced status — the leading
slash was gone, and what left the composer was a message, not a
command. completeSlashWord re-attaches exactly one slash and strips
any the name already had, so the box can never come to hold //.
The rows themselves come from flattenAvailableCommands
(webapp/lib/slash-routing.ts), which flattens the availableCommands
dict deduped on first occurrence: the mcode group (engine
commands over ACP) and the webui group both carry a help, and the
palette keys its rows by name, so an unchecked flatten rendered the same
key twice — React logged "Encountered two children with the same key"
and the user saw two identical rows. The duplicate is the same slash
command to the user (typing it routes through routeSlashInput, not the
row), so the palette shows one entry per name.
webapp/test/slash-commands.test.ts pins the dedupe against the real
function the composer calls.
| Design | Enter on an ambiguous prefix | Rejected because |
|---|---|---|
| Enter always sends (shipped) | runs /co, which the engine rejects and the user sees at once |
— |
| Enter sends unless an arrow key moved the highlight | runs the highlighted candidate | one key then means two things depending on state the user did not necessarily set — hover moves the highlight too — and the failure is silent: a command nobody typed is what runs |
| Enter always completes | inserts the first candidate | the reported defect; a typed command is unreachable without arrow keys |
The palette does not replace the send button and the send button does
not replace the palette. The composer's hint line states the two
bindings outright: Enter 发送,Tab 插入.
The response is written after the dispatch, so it describes the command rather than the receipt.
| Status | Body | Meaning |
|---|---|---|
200 |
{ok:true, cmd} |
the dispatcher claimed the command and ran it |
400 |
{ok:false, error, reason:"unknown_command", cmd, knownCommands[], suggestion} |
nobody claimed it; nothing was mutated |
4xx |
request gate refusal, before the handler runs | untrusted Origin, bad token (403), rate limit (429) |
5xx |
gate, audit, or handler failure | the write-ahead audit is fail-closed by design |
A declined authorize("slash.clear") gate is not an error status:
handleCmdCommand appends ● 已取消 /<cmd> (授权未通过: <decidedBy>)
to the transcript and still answers 200 {ok:true, cmd}, with nothing
mutated. So 200 is not proof that the command did something — read
the transcript.
error is the one-line Chinese string the composer's error banner
shows; reason is the machine-readable discriminator; suggestion is
the fix — "send it as a normal message" for a /api/send command such
as /goal, otherwise the list of commands this endpoint does accept.
knownCommands carries the accepted set so a client can render it
without hard-coding the list.
An earlier revision of this route wrote 200 {ok:true} before
dispatching, which made every input a success — /goal <text> cleared
the composer and did nothing at all.
A rejected command mutates nothing: no chat line, no goal, no session.
The composer restores the rejected text (merged after anything typed
while the request was in flight) and shows the banner; a command routed
to /api/send that the engine rejects surfaces as an error alert on the
anomaly channel.
/api/cmd output and engine output are both transcript lines, and they
do not come from the same place.
| Kind | Written by | In the engine runtime DB? | Survives a poll tick? |
|---|---|---|---|
engine turn (› ping, ● pong, tool blocks) |
the engine, streamed into cs.chat |
yes | yes, refreshed from the DB |
the streaming mirror of that turn (one folded ● answer…, a → bash header without its args) |
the same stream, into the same array | yes — the same text, folded | no — it retires, the engine's own lines take its place |
/api/cmd echo (› /help, ● 可用命令:…, ● 当前 model=…, ● 变更概览 …) |
interaction/commands.js, into cs.chat |
no — the engine never sees it | yes, and it is the only thing that keeps it there |
| a turn another client ran (desktop app, TUI) | the engine, for a different cid | yes | yes, pulled in — that is the poll's purpose |
The four-second poll (lib/transcript-sync.js, MCODE_WEBUI_TRANSCRIPT_SYNC_MS=0
disables it) re-reads the engine's view so a conversation driven elsewhere
catches up in an open tab. It is a merge, not a replacement:
mergeEngineTranscript (lib/transcript.js) walks the engine read and
the lines already shown in lockstep, keeps any line the engine does not
know about in place, and appends the engine's remainder.
The mirror is the third kind, and it is the one the merge has to retire. While
a turn streams, the same engine output is written into cs.chat a second time
in a folded form — the answer and thinking branches write one line
(prefix + text.replace(/\n+/g, " ").trim()) where the engine's own mapper
keeps one array entry per source line, and a tool header is written → bash
when the frame carried no rawInput against the engine's
→ bash {"command":…}. Neither can ever be byte-equal to what the engine
holds, so the lockstep walk called every mirror a locally-authored line, kept
it, and appended the engine's whole spine behind it. The answer, the tool block
and the thinking chain each rendered twice, and persistCurrentChat made the
duplicate permanent. Measured on a UAT session: 81 stored lines against a
67-line engine read, 14 of them a second copy of engine content.
Retirement is an identity test, not a shape heuristic. A folded prose mirror
(●/▲/›/○) retires when the maximal run of engine lines carrying the
same glyph, starting at the cursor, folds — their texts joined by a single
space, every whitespace run collapsed — to exactly the mirror's folded text. A
tool header retires when the engine line at the cursor names the same tool, and
the indented block goes with it on both sides. So the engine must already hold
that text at that position: a /api/cmd echo, which the engine has never seen,
has no fold to match and is kept. And a test that misses — unusual spacing, a
tool block the engine has not finished writing — leaves the line in place, which
is the double render the merge already had. No path drops content the engine
read did not account for.
Server-written annotations — §§ processed_duration=Nms, §§ turn_msg=<id>,
##tc:<id> — are the one class of engine line the merge may not treat as an
ordinary line, because position is their entire meaning: the decoder resolves
each one onto the block above it. A tab whose chat was recorded before its
marker shipped does not carry the line, so the merge emits it at the cursor
the engine put it at and never at the tail. Two consequences, both visible in
the chat. A turn another client ran keeps its own turn coordinate instead
of handing it to whatever the user ran next, which is the difference between
the 「已编辑 N 个文件」 card reading this turn's diff and reading the engine's
latest turn. And a transcript recorded before a marker shipped is annotated in
place rather than replayed behind its own copy — the tail position duplicated
the whole conversation instead, once annotated and once not.
The alternative that was rejected: assign the read over cs.chat. It is
one line, and it is what shipped. The consequence was that /help and
/status returned 200, cleared the composer, rendered their output for
about four seconds, and then vanished — and persistCurrentChat recorded
the deletion, so a reload did not bring them back. Measured on a live
instance: the echo was on the wire at t+200 ms and gone by the next tick.
An in-memory ledger of local lines was also considered and rejected: it
would not survive the reload it was meant to protect, which is exactly the
half of the defect users noticed.
The merge assumes the engine appends and never rewrites a line it has
already emitted. A rewrite would show up as the old line sitting next to
its replacement rather than being replaced; the switch path's backfill rule
(routes/sessions.js) already depends on the same assumption.
POST /api/send writes 200 {ok:true} at the top of handleSend and
runs the turn afterwards. The acknowledgement therefore reports receipt,
and the deadline the browser imposes on it (SEND_ACK_TIMEOUT_MS, 30 s in
webapp/lib/api.ts) reports round trip. Neither says whether the engine
took the prompt — during a stalled proxy or a busy event loop the engine can
be executing the message while the browser is still waiting.
Reporting that as a failure is a claim about a side effect that may already
have happened, and the composer's response to a "failure" — put the text
back in the box — turned it into a duplicate execution. In testing, a
sleep 35 ran twice because the first attempt's acknowledgement was slow
and the user pressed Enter again.
| Old | New | |
|---|---|---|
| Error shape | Error("no response within 30000ms"), matched by wording |
SendUnconfirmedError, matched on an unconfirmed flag (isSendUnconfirmed) |
| Deadline | 30 s | 30 s — unchanged; a longer one only moves the same false negative later |
| Decision | none — the deadline was the verdict | probeSend asks GET /api/state, bounded to 3 reads over ~2.7 s |
| Draft restored | always | only when the server positively holds no record of the send |
| Banner | 消息发送失败: no response within 30000ms, red |
one of three, none of which claims failure |
probeSend (webapp/lib/send-confirmation.ts) reduces its reads to one of
three answers:
| Answer | Evidence | Draft | Banner says |
|---|---|---|---|
accepted |
a turn is running for this cid, or the prompt's › echo is in the transcript |
not restored | sent, never confirmed, the engine is running it — do not send it again |
rejected |
the server answered and holds no record | restored | not delivered, the server has no record; the text is back in the box |
unreachable |
no read came back | restored | status unknown, it may already be running — check the history before sending again |
accepted is the case the whole design turns on: a send the engine may
already be running must never come back as text one Enter can re-send.
unreachable restores the draft even though the answer is unknown, because
losing what the user typed is the worse defect, and the banner carries the
"check the history first" instruction that makes the restore safe. The
banner is also styled as secondary text rather than as an error.
The banner's display semantics are the three answers above; its dismissal
is separate (webui-parity 106). While running.active is up, the warning is
doing its job. When the flag falls — the turn it warned about is over — the
grey banner goes with it (unconfirmedPatchOnTurnEnd in
webapp/lib/composer-draft.ts, applied by a composer effect that watches the
running-flag fall): after sleep 35 finished, the banner used to sit under
the input until the next send or a reload. A real rejected refusal keeps
its dismiss paths; no display rule changed.
The banner is also addressed, not broadcast. The draft store is keyed by session, and the catch branch writes the banner into the key of the session the send was dispatched FROM — so a failure recorded in session A while the user has already switched to session B never paints B red; the user finds the banner when they return to A. The previous behaviour (a module-scope shared box, then #141's clear-on-switch) either bled the banner across sessions or destroyed the returning session's own unread one.
A client-generated idempotency key on POST /api/send would make the
duplicate structurally impossible rather than merely unlikely. It is not
implemented: it is a request-contract change, and it needs a
server-side dedup store with a defined window. Treated as its own ticket,
not folded into this fix.
How you would tell it works. Send /help in a session that already has
an engine turn, and leave the tab open: the output is still there ten
seconds later, and it is still there after a reload. Force an
acknowledgement timeout against a server that is running the turn: the
banner says the engine is running the message, and the composer is empty.
Wait for the turn to finish: the grey banner disappears on its own.
Everything the user has parked in the composer — typed text, attachment
chips, the send-error banner — is stored under the active session's key
(webapp/lib/composer-draft.ts, a Map keyed by state.sessionId; "" is
the no-session home-screen bucket). Switching sessions swaps the whole box:
session B never shows session A's draft or banner, and both survive the
round trip. The smoke run's s28 capture was the shared-bucket version of
this store: session 2's view showing session 1's draft, 409 banner and
model chip at the same time.
Per-session storage, not clear-on-switch, is the deliberate choice: a
clear-on-switch effect (the #141 interim fix) also fires when the user
comes BACK, destroying the very draft and unread banner they returned for.
Keyed storage keeps the good half of the old global behaviour (nothing is
lost when hopping between sessions) while removing the bleed. Drafts are
not persisted to localStorage — they are working state for the current
page visit; the persisted surface stays lib/persist.ts's contract.
The model picker's chip VALUE always read the server snapshot and needs no
isolation; its local UI state (open cascade, previewed row, per-model draft
mirror) resets when the session key changes, so no menu state from session A
visually persists into session B's view. Whether a model pick made in one
session's view can land in another session's engine config is a
server-side applyConfigOptionUpdate question and out of this ticket's
frontend scope.
How you would tell it works. Type a draft in session A, switch to session B: B's composer is empty and the chip follows B's server model. Switch back: A's draft and any unread failure banner are exactly as left.
Every /api/* endpoint listed below is registered either by Hono
(packages/webui/server/app.js) or by the legacy dispatcher
(packages/webui/server/router.js); the file path is the implementation
of record. OWNED_ROUTES (Hono) is the ledger (62 routes), and the legacy
dispatcher owns the two SSE channels (/api/events, /api/alerts) plus
the static + trajectory mounts.
On /api/events the server sends named frames alongside the state
snapshot — needs_authorization, authorization_decided,
token.first_run, auth.token_rotated, providers.updated,
session-tree-changed, heartbeat. session-tree-changed (Agent Team)
carries no payload (data: {}); it fires when a subagent row lands in the
runtime db and tells the sidebar to re-fetch GET /api/session-tree, and it
bypasses the push coalescer so a sparse tree update is never dropped. The
parent's subagent list rides the state snapshot as recentSubagents[]
({toolCallId, sessionId, agentName, status, createdAtMs, updatedAtMs}),
idempotent on toolCallId, capped at 32 with a 5-minute TTL; the chat
renderer matches a → task block to its entry by toolCallId (the ##tc:
marker), not by tool name.
| Method | Path | Handler file | Notes |
|---|---|---|---|
GET |
/api/health |
routes/health.js |
200 {ok, port, defaultModel, defaultWorkspace, mcodeCmd, mcodeVersion, maxConcurrent} |
GET |
/api/account |
routes/account.js |
200 engine-projected card; {ok:false, reason:"no_client"|"rpc_error"|"account_unavailable"} when the engine hasn't answered |
GET |
/api/state |
routes/state.js |
full state projection (snapshot) |
GET |
/api/sessions |
routes/sessions.js#handleListSessions |
merged webui + mcode session list |
POST |
/api/sessions |
routes/sessions.js#handleNewSession |
creates a webui session record |
POST |
/api/sessions/switch |
routes/sessions.js#handleSwitchSession |
swaps the active conversation |
POST |
/api/sessions/rename |
routes/sessions.js#handleRenameSession |
retitles a session (B03 authorize-gated) |
GET |
/api/sessions/search |
routes/sessions.js#handleSearchSessions |
cross-workspace fuzzy search (B03 authorize-gated) |
POST |
/api/sessions/cleanup-orphans |
routes/sessions.js#handleCleanupOrphans |
drops mcode sessions no webui record references (scope=orphans|all) |
DELETE |
/api/sessions/:id |
routes/sessions.js#handleDeleteSession |
B03 authorize-gated; deletes from webui + mcode sqlite |
GET |
/api/session-tree |
routes/sessions.js#handleSessionTree |
sidebar tree projection |
GET |
/api/acp-sessions |
routes/sessions.js#handleAcpSessions |
mcode acp session list |
GET |
/api/acp-session-title |
routes/sessions.js#handleAcpSessionTitle |
title helper for ?sid=... |
GET |
/api/sessions/:id/export |
routes/export.js |
?format=md|json[&download=true]; 400 on bad format; 403 on authorize decline; 404 on missing session |
POST |
/api/send |
routes/chat.js#handleSend |
fire-and-forget; 200 {ok}; 400 content required; 409 {reason:"cid-busy"|"session-busy"|"at-capacity", running?, limit?}. The claim is per conversation, so a second conversation of the same tab is not blocked — see Parallel turns in one tab; the idle watchdog aborts a run that stays silent for MCODE_WEBUI_PROMPT_IDLE_TIMEOUT (default 120 s) |
POST |
/api/stop |
routes/chat.js#handleStop |
200 {ok, wasRunning, cancelled, hardKilled, note} |
POST |
/api/cmd |
routes/chat.js#handleCmd |
the eight button commands only; 200 {ok, cmd} when claimed, 400 {ok:false, reason:"unknown_command", knownCommands, suggestion} when not — see Slash commands |
POST |
/api/usage |
routes/usage.js#handleUsage |
record-only + projection |
POST |
/api/usage-trigger |
routes/usage.js#handleUsage |
alias kept for legacy clients |
GET |
/api/usage-real |
routes/usage.js#handleUsageReal |
real-token snapshot |
POST |
/api/refresh |
routes/usage.js#handleRefresh |
force a refresh |
GET |
/api/usage/forecast |
routes/usage.js#handleForecast |
linear + Huber extrapolation of exhaustion time |
POST |
/api/workspace |
routes/workspace.js#handleWorkspace |
{ok, error} of the workspace change |
GET |
/api/workspace/browse |
routes/workspace.js#handleWorkspaceBrowse |
?path=<abs>; 400 on bad path |
GET |
/api/workspace/tree |
routes/workspace.js#handleWorkspaceTree |
full workspace → sessions tree |
GET |
/api/workspace/resolve |
routes/workspace.js#handleWorkspaceResolve |
?name=<folder> → candidate absolute paths |
GET |
/api/workspace/recent |
routes/workspace.js#handleWorkspaceRecent |
?search=&limit= (limit clamped ≤ 20) |
GET |
/api/fs/read |
routes/fs.js#handleFsRead |
?path=&showHidden=1; containment gate; 400 missing path |
GET |
/api/fs/read-file |
routes/fs.js#handleFsReadFile |
?path=&confirm=1; 200 text/JSON (success carries mtime, the conflict-detection baseline for POST /api/fs/write); 403 {code:"credential"} on a credential shape (unless confirm=1); 413 oversize (fs-util DEFAULT_FILE_READ_MAX = 512 KiB); 415 binary / non-regular |
GET |
/api/fs/raw |
routes/fs.js#rawStreamToWebResponse |
?path=&download=1&confirm=1; streaming 20 MiB cap; same credential gate; mime-by-extension table including .html/.htm, .svg, .png/.jpg/.gif/.webp, .js/.mjs/.css/.json |
POST |
/api/fs/mkdir |
routes/fs.js#handleFsMkdir |
{path}; parent in allowed roots; 403 on containment fail |
POST |
/api/fs/write |
routes/fs.js#handleFsWrite |
{path, content, expectedMtime?, expectedSize?, confirm?} — the preview editor's save (slice 27). 200 {ok, path, size, mtime} (fresh baseline); 400 {code:"missing-path"|"missing-content"|"invalid-content"|"not-a-regular-file"}; 403 containment / 403 {code:"credential", credentialReason} (slice-16 shapes without confirm:true); 404 {code:"not-found"} (vanished file — TOCTOU guard; a missing path normally fails the gate first, same as reads); 409 {code:"conflict", diskMtime, diskSize} (stale baseline, nothing written); 413 {code:"too-large"} (write cap = the read's 512 KiB). Bare writeFileSync on the gated path — no shell anywhere. confirm:true on a credential shape emits the credential.override audit line with endpoint:"write". |
POST |
/api/fs/open-default |
routes/fs.js#handleFsOpenDefault |
{path}; 400 {code:"missing-path"} / 403 {code:"out-of-bounds"} / 400 {code:"not-a-regular-file"} / 503 {code:"no-opener"} / 502 {code:"spawn-failed"} |
POST |
/api/fs/reveal |
routes/fs.js#handleFsReveal |
{path}; same code → status map as open-default. Consumers: the file-preview toolbar and the sidebar project menu's 在文件夹中显示 (SB-6) |
GET |
/api/fs/search |
routes/fs.js#handleFsSearch |
?root=&q=&depth=&maxNodes=&wallMs=&limit=&includeHidden=1; 400 {code:"missing-root"|"missing-q"|"not-a-directory"|"stat-failed"}; success envelope: {ok, root, q, matches:[{path,name,type,ancestors,credential?,credentialReason?}], scanned:{dirs,files,total}, skipped:{node_modules,n,.git,n,credential,n,huge,n,optional:{dist,build,…}}, truncated, truncatedReason: null|"depth"|"nodes"|"wallClock"|"matches", elapsedMs, budgets}. Walker defaults: maxDepth=8, maxNodes=5000, wallMs=1500, maxMatches=200; absolute limits: 16/50_000/5_000/1_000 (packages/webui/server/lib/fs-search.js). node_modules and .git are non-overridable skips. |
GET |
/api/git/status |
routes/git.js#handleGitStatus |
?dir=; 400 {error:"missing dir"} |
GET |
/api/git/branches |
routes/git.js#handleGitBranches |
?dir=, leading * → current flag |
GET |
/api/git/diff |
routes/git.js#handleGitDiff |
?dir=&file=; 400 {error:"missing dir/file"}; falls back to --no-index for untracked |
POST |
/api/git/checkout |
routes/git.js#handleGitCheckout |
{dir, branch}; branch allow-list ^[A-Za-z0-9._/-]+$ + leading-dash guard; 400 {error:"missing dir/branch"}; 413 {code:"BODY_TOO_LARGE"} on cap |
GET |
/api/settings |
routes/settings.js#handleGetSettings |
full settings projection |
POST |
/api/settings |
routes/settings.js#handlePostSettings |
500 {error:"audit write failed"} if event log fails; B03 authorize gates within the handler |
POST |
/api/auth/decision |
lib/authorize.js#handleAuthDecision |
{requestId, approve}; 200 resolved; 404 no such pending request; 400 bad body; idempotency guard via resolved-set delete |
POST |
/api/upload |
routes/upload.js |
multipart required; 400 if not; 413 {code:"UPLOAD_REQ_TOO_LARGE"|"UPLOAD_FILE_TOO_LARGE"|"UPLOAD_QUOTA_EXCEEDED"}; 400 {code:"UPLOAD_MALFORMED"|"UPLOAD_ABORTED"}; write-ahead audit upload.create.intent before disk, upload.create after; 200 {ok, path, name, size} |
GET |
/api/models |
routes/model.js#handleGetModels |
engine model + webui label/limit projection; thinkingLevels from both engine thinking schemas (effort list verbatim, switchable builtins as ["off","on"]). Response {ok, models, groups, current, currentThinking, source, reason?} — models the flat list; groups provider-grouped for the picker ({id, label, auth:{hasKey,type}, protocol?, models}, auth/protocol only on config groups — __engine/minimax_api carry id/label/models); current the active id or null (never a fabricated default); currentThinking the active level (thinkingEffort.currentValue → cs.model.thinking → null); source = acp-session-config|config+mcode-cli-bundle|mcode-cli-bundle (which layer answered); reason:"no_catalogue" only when models is empty |
POST |
/api/set-model |
routes/model.js#handleSetModel |
{model, thinking?}; 400 only when model is empty and thinking is absent (missing-parameter, not unknown-model — an unknown model name is recorded and pushed, never validated here); effort models push model+thinkingEffort, variant models fold the on/off level into one model selection. Gated on authCredentials.setThinkingEffort for a standalone effort write only (M3-B14): a pure model switch and a variant-channel pick are not gated |
POST |
/api/permissions |
routes/model.js#handleSetPermissions |
{mode}; mapped to engine mode via WEBUI_TO_MCODE_PERMISSION; gated on authCredentials.setPermissionMode, which is behaviourally inert today (M3-B14) |
GET |
/api/permissions-modes |
routes/model.js#handleListPermissionModes |
engine's current availableModes |
POST |
/api/answer |
routes/model.js#handleAnswer |
Removed capability — tombstone only. Always 410 {ok:false, removed:true, error}. It used to answer 200 {ok:true, deprecated:true} without reaching the engine, and four buttons called it, so a click looked successful while the prompt stayed pending. webapp/lib/api.ts deliberately exports no client for it; do not add one without a channel that reaches the engine. See "Blocking prompts: what each one can actually answer" |
GET |
/api/providers |
routes/providers.js#handleGetProviders |
masked catalogue |
PUT |
/api/providers |
routes/providers.js#handlePutProviders |
full replace; 400 on validate, 500 on write failure |
POST |
/api/providers/test |
routes/providers.js#handleTestProvider |
{provider}; structured codes → status |
GET |
/api/providers/presets |
routes/providers.js#handleGetPresets |
gallery |
POST |
/api/providers/preset/:id/enable |
routes/providers.js#handleEnablePreset |
one-click enable |
GET |
/api/model-source |
routes/model-source.js#handleGetModelSource |
{ok, source, apiKey:{available,hasKey,masked,testState,lastTestedAtMs}}; 501 when the host has no getMiniMaxModelSource, 503 when no runtime is booted, 502 {code:"UNKNOWN_MODEL_SOURCE"} for a value outside the engine's own two. available:false is not hasKey:false |
PUT |
/api/model-source |
routes/model-source.js#handleSetModelSource |
{source}; 400 {code:"INVALID_MODEL_SOURCE"|"BAD_FIELD_TYPE"}, 400 {code:"NO_API_KEY"} when the engine refuses the BYOK direction; the response carries what the engine PERSISTED |
PUT |
/api/model-source/api-key |
routes/model-source.js#handlePutModelSourceApiKey |
{apiKey, saveAndUse?}; an absent/empty/whitespace apiKey is the KEEP sentinel → 200 {changed:false} with no engine write; 400 {code:"BAD_FIELD_TYPE"|"INVALID_API_KEY"}; 500 {code:"engine_error"} never carries the thrown message |
POST |
/api/model-source/test |
routes/model-source.js#handleTestModelSource |
{modelId?}; always 200 for a COMPLETED probe ({ok, success, providerId:"minimax_api", tested:"stored_key", status}) including success:false; non-200 only when the probe is refused (503/501, or the engine's 400 NO_API_KEY) |
POST |
/api/follow-up |
routes/follow-up.js#handleFollowUp |
{behavior:"queue"|"steer", content, attachments?, requestId?} — the engine session id comes from the server's own conversation state, never the body; 400 {code:"invalid_follow_up_behavior"|"follow_up_empty"|"no_active_conversation"|"BAD_FIELD_TYPE"}; 409 {code:"no_active_turn"|"turn_not_owned"} when this process does not own the running turn, and nothing is queued; 501 when the host lacks the method, 503 when no runtime is booted; 200 {ok, behavior, itemId, position, status} (queue) or {ok, behavior, turnId, mode} (steer) — the engine's own answer |
GET |
/api/logs/export |
routes/logs.js#handleExportLogs |
always 200 text/plain with Content-Disposition: attachment; filename="mcode-webui-logs-<YYYYMMDDTHHMMSS>.txt" — the crash trail (WEBUI_DATA_DIR/.server.err) plus the last 2000 lines / 2 MiB of the event log, each section stating its own truncation or absence. No status code for a missing or unreadable log file: that is reported inside the body, because the client is a browser anchor and a 4xx would be saved as a .txt file containing JSON |
POST |
/api/debug/inject |
routes/debug.js#handleDebugInject |
DEBUG_INJECT=1 gate |
GET |
/api/debug/state |
routes/debug.js#handleDebugState |
same gate |
POST |
/api/protocol/set-mode |
routes/protocol.js#handleSetMode |
mid-session mode change |
POST |
/api/protocol/set-config-option |
routes/protocol.js#handleSetConfigOption |
with configId:'permissionMode' this becomes the mid-session mode switch |
POST |
/api/protocol/cancel |
routes/protocol.js#handleCancel |
acp session/cancel notification |
POST |
/api/protocol/load-session |
routes/protocol.js#handleLoadSession |
?cwd=, fallback to current |
POST |
/api/protocol/activate-session |
routes/protocol.js#handleActivateSession |
one acp client tracks one active session |
GET |
/api/protocol/list-sessions |
routes/protocol.js#handleListSessions |
?cwd= filtered |
GET |
/api/protocol/capabilities |
routes/protocol.js#handleCapabilities |
{mcodeVersion, mcodeName?, mcodeTitle?, capabilities, capabilitiesProvider, capabilitiesProviderFor, capabilitiesUnavailable, notes} — capabilities is the engine's declared 14-key capability object (it was the ACP wire table MCODE_ACP_CAPABILITIES before M3-B4) |
| Method | Path | Reason it stays here |
|---|---|---|
GET |
/api/events |
SSE channel: writer held by lib/state-bus.js across pushes (Hono streaming lands in P2) |
GET |
/api/alerts |
Independent anomaly SSE channel (bell icon + audit log) |
GET |
/trajectory, /trajectory/... |
Separate panel; SPA fallback to /trajectory/ |
GET |
/ and /index.html |
serveIndex / auth-gate.html |
GET |
*.<ext> |
static (webapp/out) |
OPTIONS |
* |
204 short-circuit (CORS preflight) |
The Hono layer still owns /api/health and /api/settings for every real
consumer; the legacy copies exist only so gate tests
(checks/router-origin-gate.check.mjs) have a route to assert against.
Every HTTP request that crosses a destructive boundary waits on a per-cid authorize round-trip (5-minute default timeout, fail-closed):
session.delete—DELETE /api/sessions/:idsessions.cleanup-orphanssession.cleanup-all(extension hook)session.export—GET /api/sessions/:id/exportsession.search—GET /api/sessions/searchtoken.resetslash.clear—/clearand/newon the chat streamstartup.cleanup— boot-time orphan sweep
The whitelist is the only source of truth — anything not on this list cannot be gated via the modal flow.
The 5-minute budget answers "a human saw the modal and did not answer".
It does not answer "no human was ever shown one". The gate is a push:
pushAuthRequest writes a needs_authorization frame into the requesting
tab's SSE response, and the decision comes back on POST /api/auth/decision. When the request's client has no live connection
(state-bus.js#hasDecisionListener — no response registered for that cid,
or the registered one can no longer be written to), nobody can decide, so
the fail-closed result is already determined. The request is answered at
once with the ordinary decline body — 403 {ok:false, error:"authorize declined", decidedBy:"timeout", decidedAt} — and audited as
auth.unreachable with reason:"no_connected_client", so an operator can
tell "nobody was there" from "somebody said no". Before this the same call
held the socket open for the full five minutes with no status and no body,
which is indistinguishable from a hang; in practice that is what
curl -X DELETE /api/sessions/<id> from a script saw, because a request
without ?cid= has no client to ask.
A bus that cannot answer the question (a test double that does not model the connection registry) is treated as "might have a listener" and keeps the old wait. The short-circuit can only ever deny — no path approves anything without a recorded decision.
"Is anybody listening?" and "does this request have an owner?" are two
different questions, and the rule above answers only the first. A
request without ?cid= — a curl, a script, a caller that forgot
withClientQuery — is told "yes, somebody is listening" as soon as one
browser tab is open, because an empty cid is the broadcast target and a
connected tab really can see and answer the modal. Nobody asked for it,
so nobody answers it, and the destructive request sits there for the full
300000 ms. A gate that fires only while the connection registry is empty
therefore misses the common case, and the caller sees a hang rather than
a denial.
The routes that serve an identified HTTP caller pass
{requireRequester: true} for that reason: session.delete,
sessions.cleanup-orphans, session.export, session.search. With no
owner the answer is the same fail-closed one, at once, audited as
auth.unreachable with reason:"no_requester" so an operator can tell it
apart from a closed tab. The rule is opt-in because the difference is
load-bearing in the other direction too — startup.cleanup asks with an
empty cid on purpose, and any tab may answer it.
DELETE /api/sessions/:id applies it before anything else, ahead of the
plan, so an unattributable delete costs no store read and reaches no
engine at all. The same handler also stops asking a question it can
already answer: an id that is absent from the session store and is not
an mvs_ sid has no wrapper to splice and no engine rows to remove, so
it returns the 404 {ok:false, error:"session not found"} this branch has
always returned — the facade's own not_mcode_sid / already_absent
pair, stated at the HTTP layer instead of waited out — with no governance
round-trip. Every id that can delete something, a resolved record or an
orphan mvs_ sid whose engine rows are about to go, still passes the
gate.
components/modals.tsx renders three blocking prompts. Two of them carry a
decision the engine receives; one cannot, and says so instead of pretending.
The distinction is a contract, not a UI preference: a button that posts
somewhere the engine never reads accepts the click and leaves the prompt
pending, which is worse than showing no button.
| Prompt | Answer channel | Engine receives it? |
|---|---|---|
| ask_user | POST /api/send {content, isAskAnswer:true} |
Yes. routes/chat.js reads isAskAnswer and forwards the string; options, free text and Skip all use it. |
| Authorization | POST /api/auth/decision {requestId, approve} |
Yes — for webui's own authorize() actions. It is a webui-local gate, not the engine's tool-permission prompt. |
| Plan review | none | No. The prompt is read-only and renders no decision. |
A plan review is not an ACP message. It is a runtime questionnaire:
local-runtime-v2opens it asquestionnaire.askwithmode:'plan'and a single step carrying one option,approve(packages/local-runtime-v2/src/service/plan/application.ts:278).- The ACP bridge projects it one way, as a
plan_updatenotification (packages/tui/src/acp/agent.ts:1356). Nothing carries the answer back. - The TUI answers it on the local-runtime channel —
runtime.replyQuestionnaire(packages/tui/src/tui/controller/interaction/interaction-flow.ts:938) — which this package does not speak. - The fallback is an incoming
session/requestPermissionrequest (packages/tui/src/acp/interactions.ts:680).acp.mjs#_dispatchemits it with no responder, so it cannot be answered either. The agent's entireapp.onRequest(acp.methods.agent.*)surface isinitialize,authenticate,session.new/list/fork/load/resume/close/setMode/setConfigOption/prompt— there is no plan-decision method to call.
So the modal shows the plan document and states that the review must be answered elsewhere. It is dismissible: with no working button, a dialog that cannot be closed is a trap. Closing it does not answer the review, and the turn stays pending on the engine either way.
The only producer is agent.ts:1356, and it sends
{sessionUpdate:'plan_update', plan:{type:'markdown', planId, content}}. The
projection in server/lib/mcode-acp.js reads exactly that. It used to read
planId / title / summary / options off the top level of the update,
where the engine puts none of them — so plan.active was true with an empty
title, an empty body and no options. options is always empty and stays in the
type only so a consumer cannot trip over undefined; the review's single
approve option lives on the questionnaire side.
Two things must land together, in this order:
- The ACP client must answer incoming requests.
acp.mjs#_dispatchcurrently drops any message that is a request rather than a response. It needs a responder registry keyed by JSON-RPC id, or the engine'ssession/requestPermissionhangs until its projection times out and fail-closes by dismissing the questionnaire. - Only then fix
initialize.acp.mjs#startsendscapabilities: {mcpCapabilities: …}, but the agent readsparams.clientCapabilities(agent.ts:434), so webui negotiates no client capabilities at all. The consequence is larger than the plan: theplan_updateprojection is gated onclientCapabilities.planand therefore never fires today, and the elicitation path that would answer a multi-option questionnaire is unavailable for the same reason. Fixing the field name alone would start sending questionnaires webui cannot answer.
POST /api/answer stays a 410 tombstone until then, and
webapp/lib/api.ts exports no client for it.
See packages/webui/docs/ARCHITECTURE.md for the runtime topology, request lifecycle, and SSE contract. In short: packages/webui/server.js registers the workspace import resolver and delegates to server/bootstrap.js; server/router.js applies the gate chain (CORS → origin/CSRF → LAN → token → rate limit → read-only) and dispatches to server/routes/*; server/lib/* holds one-concern modules; acp.mjs is the ACP client spawning the engine; webapp/out/ (the Next static export) is the UI, with public/trajectory/ and public/auth-gate.html (served from the export root) as the only remaining legacy assets.
The HTTP layer is split: legacy routes cover the streaming/SSE surfaces and the auth-gate/html fallback; everything else is on a Hono app (server/app.js). OWNED_ROUTES is the literal-source greppable ledger; ownsRequest(method, pathname) resolves the runtime dispatch decision against the Hono router table.
server/trajectory/ (migrated from the mcode-trajectory-studio plugin) inspects local sessions read-only via the runtime SQLite projection with messages.jsonl fallback, offering turn/duration/token/compaction/subagent views. It is mounted at /trajectory/ behind the webui's gates and can also run standalone:
node packages/webui/server/trajectory/main.mjs --serve # loopback panel
node packages/webui/server/trajectory/main.mjs --doctor # data-source diagnostics
node packages/webui/server/trajectory/main.mjs --stdio # MCP over stdio (7 tools)The seven MCP tools are registered in packages/webui/server/trajectory/mcp.mjs:
trajectory_list, trajectory_summary, trajectory_get, trajectory_search,
trajectory_tasks, trajectory_task_output, trajectory_studio. Server
name: mcode-trajectory-studio, version 0.1.1. Protocols supported:
2025-06-18, 2025-03-26, 2024-11-05 (newest first).
pnpm --filter @mavis/webui test # full node:test suite (unit + mocked + integration + matrix + trajectory)
pnpm test:webui # same, from the repository root (CI gate)
pnpm test:webapp # webapp (browser) suite: render + interaction tests
node packages/webui/scripts/check-docs-alignment.mjsThe package has three runtime dependencies (hono + @hono/node-server for the HTTP layer, @mavis/shared for the workspace path contract) and requires Node 22.19+ (the trajectory studio additionally needs node:sqlite, floor 22.13).
pnpm test:webapp runs webapp/test/**/*.test.ts on node --test. It has two rendering tools, and they answer different questions.
| Tool | Answers | Cannot answer |
|---|---|---|
renderToStaticMarkup (react-dom/server) |
What does this page print? | Anything requiring an event, an effect or a re-render |
the DOM harness (webapp/test/helpers/dom.ts) |
What happens when the user presses a key? | Nothing about a tree that is not mounted — it is not a snapshot tool |
A defect that only appears once a keydown reaches a handler is invisible to static markup: a string has no listeners. That is not hypothetical. The Shortcuts page's capture → verdict → conflict-report path had no test at all, and the mutation that swallowed the conflict report stayed green. Keep static markup for "what does it print" (settings-extra-pages.test.ts) and reach for the harness when the answer is "what happens when".
// webapp/test/helpers/dom.ts — mount a component, drive it, unmount it.
import { withDom, mount, resetStorage } from "./helpers/dom";
test("the conflict is reported on the row that was edited", async () => {
await withDom(createElement(ShortcutsSection, { t }), async (view) => {
await view.pressKey("settings-shortcuts-binding-global-search", {
key: "O", ctrlKey: true, altKey: true,
});
assert.match(view.text("settings-shortcuts-conflict-global-search") ?? "", /新建无项目任务/);
});
});The handle mount / withDom returns:
| Member | Purpose |
|---|---|
find(id) / query(id) / findAll(id) / text(id) / has(id) |
Look up by data-testid; find throws and lists the testids that were rendered |
pressKey(target, {key, ctrlKey, altKey, shiftKey, metaKey}) |
Dispatch a bubbling, cancelable keydown and flush React |
keyEvent(press) |
Build that keydown without dispatching, to assert on defaultPrevented afterwards |
click(target) / fire(target, type, init) / type(target, value) |
The other events, all flushed |
run(fn) |
Run an arbitrary block inside act, for a raw dispatchEvent |
rerender(node) / flush() / html() |
Re-render, drain timers, serialize |
unmount() |
Detach the root; withDom does it on the throw path too |
window |
The happy-dom window, for its localStorage |
Four rules, each one a way the harness would otherwise lie:
- Import
helpers/dombefore any component import.react-domcapturescanUseDOMwhen it is first evaluated; with no window at that moment it falls back to a host config with no event system. The harness publishes the window in its own module body and pullsreact-dom/clientin dynamically, so the order is safe as long as the harness import comes first. Getting it wrong is loud rather than silent: the tree renders empty and the first assertion fails. - Targets are
data-testidstrings, matching the convention the components already use. A missing one throws with the testids that are on screen. resetStorage()between tests. One window, one storage origin, exactly like a browser tab — state you did not clear leaks into the next test.webapp/test/helpers/dom-shim.tsis still the right tool for the markdown walker. It serves aDOMParseroverparse5and needs no window at all.
The webapp suite runs on node:test, not Vitest's DOM environment, so @testing-library would bring a beforeEach / auto-cleanup protocol this runner does not have — and the three APIs it would add are the three the handle above already exposes.
| Candidate | Transitive deps | Cost | Why not |
|---|---|---|---|
jsdom 30 |
22 | undici + css-tree + whatwg-url, ~20 MB |
Reference-complete. The suite asserts on attributes, text and event delivery — none of which is where the two implementations diverge in practice. |
happy-dom 20 |
4 (entities, whatwg-mimetype, buffer-image-size, ws) |
8 MB unpacked | Chosen. |
| neither | 0 | — | The parse5 shim in dom-shim.ts shows the middle path works for a parser, but a DOMParser cannot dispatch an event. |
happy-dom is a @mavis/webui devDependency, so it never reaches the product bundle. It does widen pnpm-lock.yaml: vitest declares it as an optional peer, so the vitest resolution key changes in every workspace importer. Any checkout that consumes the lockfile needs pnpm install --frozen-lockfile afterwards. Measured on this machine: 103 ms to import and 3 ms to construct a window, paid once per test file, and only by files that import the harness.
The package migrates the community mcode-webui plugin (v1.0.0 → v2.0.0, MiniMax-Code-Plugins PRs #16/#23/#31/#55) and the mcode-trajectory-studio plugin (PR #56) into the product. The full people and history record is co-builders.md.