diff --git a/docs/webui.md b/docs/webui.md index 59a45d87..d86fbb85 100644 --- a/docs/webui.md +++ b/docs/webui.md @@ -240,6 +240,7 @@ Default provider is `local-runtime-v2` (the only registered host provider until ### Migration state and constraints - **M1 done in this batch**: host construction (`createCatalogueHost`) moved verbatim into `server/engine/providers/local-runtime-v2.js`; `runtime-host.js` re-exports it, so every existing importer is untouched. No existing route's behaviour changed; `GET /api/engine-capabilities` is a new, additive endpoint. +- **M2 done (declaration-vs-implementation snapshot)**: `packages/webui/test/lib/engine/capability-snapshot.test.js` boots a REAL catalogue host on an isolated tmp data dir (`MINIMAX_DATA_DIR` + every `MCODE_WEBUI_*` path pinned before the provider import) and audits every `full`/`partial` key of both providers — `full` requires every tracked method to exist on the declared surface (`adapter` / `cliService` / `applications.session.diff`), `partial` requires the present half to exist, the method-named `missing` items to be genuinely absent, and kebab-case sub-capabilities (`file-write`, `git-diff`) to have no covering method; `none` is not method-checked. The tracked-method table was reflected off the live surfaces (91 adapter / 94 CliService methods), not copied from the design matrix; mutation tests in the same file pin that flipping a level, deleting a method, or growing a sub-capability each goes red. A registry-driven guard (`engine/index.js#listEngineProviderIds`) rejects any provider declaration carrying keys outside the 14-key contract, so a typo cannot pass silently. - **Capability probing (design §2.3 step 2) is deliberately not in this batch**: no route consumes a probe result yet, and wiring one would touch the catalogue host lifecycle that M1 leaves alone. It lands with the first A-batch route that needs it. - **New-provider admission rules** (enforced by the snapshot tests in `packages/webui/test/lib/engine/capabilities.test.js`): all 14 keys declared; `partial` enumerates `missing` + `reason`; declaration levels are pinned — a level flip without re-auditing the surface goes red in CI; calling an undeclared capability answers the structured 501, never an empty implementation. @@ -2039,6 +2040,17 @@ failure mode we care about is the `app/global-error.tsx` crash, not a quota error here. Per-session scroll keys are deliberate: a refresh restores the user's place in each conversation independently. +One timing invariant guards all of it (webui-parity 106): the page root +never reads these keys during render. The prerendered server HTML and the +client's first (hydration) render must be identical, and a render-phase +storage read breaks that equality the moment the `state === null` skeleton +changes shape. `app/page.tsx` renders its first frame from the shared +DEFAULT constants and applies the stored payload in one post-mount effect; +the three write-back mirrors are gated on that restore having run, so the +defaults-seeded first render cannot overwrite the stored payload. What the +user sees is unchanged: the skeleton is still up while the restore lands, +and by the time the first snapshot arrives the saved layout is in place. + ## Slash commands: which endpoint answers them (webui-parity ticket 65) A `/`-prefixed line in the composer is not automatically a command. Two @@ -2251,6 +2263,23 @@ losing what the user typed is the worse defect, and the banner carries the "check the history first" instruction that makes the restore safe. The banner is also styled as secondary text rather than as an error. +The banner's *display* semantics are the three answers above; its *dismissal* +is separate (webui-parity 106). While `running.active` is up, the warning is +doing its job. When the flag falls — the turn it warned about is over — the +grey banner goes with it (`unconfirmedPatchOnTurnEnd` in +`webapp/lib/composer-draft.ts`, applied by a composer effect that watches the +running-flag fall): after `sleep 35` finished, the banner used to sit under +the input until the next send or a reload. A real `rejected` refusal keeps +its dismiss paths; no display rule changed. + +The banner is also addressed, not broadcast. The draft store is keyed by +session, and the catch branch writes the banner into the key of the session +the send was dispatched FROM — so a failure recorded in session A while the +user has already switched to session B never paints B red; the user finds +the banner when they return to A. The previous behaviour (a module-scope +shared box, then #141's clear-on-switch) either bled the banner across +sessions or destroyed the returning session's own unread one. + A client-generated idempotency key on `POST /api/send` would make the duplicate structurally impossible rather than merely unlikely. It is not implemented: it is a request-contract change, and it needs a @@ -2262,6 +2291,38 @@ an engine turn, and leave the tab open: the output is still there ten seconds later, and it is still there after a reload. Force an acknowledgement timeout against a server that is running the turn: the banner says the engine is running the message, and the composer is empty. +Wait for the turn to finish: the grey banner disappears on its own. + +## The composer's state is per-session (webui-parity 106) + +Everything the user has parked in the composer — typed text, attachment +chips, the send-error banner — is stored under the active session's key +(`webapp/lib/composer-draft.ts`, a `Map` keyed by `state.sessionId`; `""` is +the no-session home-screen bucket). Switching sessions swaps the whole box: +session B never shows session A's draft or banner, and both survive the +round trip. The smoke run's s28 capture was the shared-bucket version of +this store: session 2's view showing session 1's draft, 409 banner and +model chip at the same time. + +Per-session storage, not clear-on-switch, is the deliberate choice: a +clear-on-switch effect (the #141 interim fix) also fires when the user +comes BACK, destroying the very draft and unread banner they returned for. +Keyed storage keeps the good half of the old global behaviour (nothing is +lost when hopping between sessions) while removing the bleed. Drafts are +not persisted to `localStorage` — they are working state for the current +page visit; the persisted surface stays `lib/persist.ts`'s contract. + +The model picker's chip VALUE always read the server snapshot and needs no +isolation; its local UI state (open cascade, previewed row, per-model draft +mirror) resets when the session key changes, so no menu state from session A +visually persists into session B's view. Whether a model pick made in one +session's view can land in another session's engine config is a +server-side `applyConfigOptionUpdate` question and out of this ticket's +frontend scope. + +**How you would tell it works.** Type a draft in session A, switch to +session B: B's composer is empty and the chip follows B's server model. +Switch back: A's draft and any unread failure banner are exactly as left. ## Endpoint catalog (against current source) diff --git a/docs/webui.zh-CN.md b/docs/webui.zh-CN.md index 8e37b6fe..a37229b7 100644 --- a/docs/webui.zh-CN.md +++ b/docs/webui.zh-CN.md @@ -240,6 +240,7 @@ GET /api/engine-capabilities[?provider=] ### 迁移状态与边界 - **本批只做迁移第一步 M1**:host 构造(`createCatalogueHost`)原样移入 `engine/providers/local-runtime-v2.js`,`runtime-host.js` 转发导出,既有引用方零改动;没有任何现有路由行为变化,`GET /api/engine-capabilities` 是纯新增端点。 +- **M2 已做(声明与实现的快照校验)**:`packages/webui/test/lib/engine/capability-snapshot.test.js` 在隔离的临时数据目录上起**真实** catalogue host(`MINIMAX_DATA_DIR` 与全部 `MCODE_WEBUI_*` 路径在 provider import 前钉死),审计两个 provider 的每个 `full`/`partial` 键——`full` 要求跟踪的方法在声明的 surface(`adapter` / `cliService` / `applications.session.diff`)上全部存在;`partial` 要求存在的部分在、方法名形态的 `missing` 项真的不存在、kebab-case 子能力(`file-write`、`git-diff`)没有覆盖方法;`none` 不做方法校验。方法跟踪表是对真实 surface 的反射取证(adapter 91 个 / CliService 94 个方法),不是抄设计矩阵;同文件的变异测试钉住改档位、删方法、子能力长出方法各自必然红。注册表驱动的守卫(`engine/index.js#listEngineProviderIds`)拒绝任何携带 14 键契约之外键的 provider 声明,拼错无法静默通过。 - **启动只读探测(设计稿 §2.3 第 2 步)本批刻意不做**:尚无路由消费探测结果,而接探测要动 M1 明确不动的 catalogue host 生命周期;随第一个需要它的 A 批路由一起落。 - **新 provider 准入规则**(由 `packages/webui/test/lib/engine/capabilities.test.js` 快照测试钉住):14 键全声明;`partial` 必须枚举 `missing` 与 `reason`;声明档位被测试钉死——不经重新审计改档位,CI 直接红;调未声明能力一律答结构化 501,绝不给空实现。 @@ -1480,6 +1481,14 @@ loading-states 相同:让 SSR 渲染测试可以脱离 `chat.tsx` 的 `@/` 别 错误。会话内每个 sessionId 单独存储滚动位置 —— 按会话恢复滚动位置 是有意为之的契约。 +一条时序不变式守着这一切(webui-parity 106):页面根组件**绝不在渲染期读 +这些键**。预渲染的服务端 HTML 与客户端首次(hydration)渲染必须逐字节 +一致,渲染期读存储会在 `state === null` 骨架屏第一次改形时炸出不一致。 +`app/page.tsx` 首帧用共享的 DEFAULT 常量渲染,挂载后的一个 effect 统一 +套用存储值;三处写回镜像都加闸在该恢复之后,默认值首帧不可能覆盖存储 +payload。用户看到的东西不变:恢复落地时骨架屏仍亮着,第一份快照到达时 +保存过的布局已经就位。 + ## 斜杠命令走哪个端点(webui-parity ticket 65) 输入框里以 `/` 开头的一行**不等于**命令。两个端点都能消费斜杠输入, @@ -1658,13 +1667,53 @@ composer 实际调用的那个函数。 回填草稿——让用户输入的内容消失是更严重的缺陷,而文案里带着「先查历史」 这句指引,回填才是安全的。该错误条同时改用次要文字色,不再是错误红。 +上面三种答案定的是错误条**何时显示**;**何时消失**是另一件事 +(webui-parity 106)。`running.active` 亮着时,灰条在履行职责;这个标志 +落下——它警告的那个回合结束了——灰条随之消失 +(`webapp/lib/composer-draft.ts#unconfirmedPatchOnTurnEnd`,composer 里 +一个盯 running 下降沿的 effect 负责套用):此前 `sleep 35` 跑完后,灰条 +会一直挂在输入框下直到下次发送或刷新。真正的 `rejected` 拒绝保持原有的 +消失路径;显示判定一字未动。 + +错误条还是**有归属**的,不是广播。草稿存储按会话分键,catch 分支把红条 +写进**发起发送的那个会话**的键下——用户在会话 A 发送失败后已经切到 +会话 B,B 的输入框永远不会因此变红;回到 A 时才看到这条失败。旧行为 +(模块级共享桶,再到 #141 的切换即清)要么把红条串到别的会话,要么把 +用户正要回去看的那个会话自己的红条销毁掉。 + 给 `POST /api/send` 加一个客户端生成的幂等键,可以让重复执行从「不太可能」 变成「结构上不可能」。本次没做:那是请求契约变更,还需要服务端带明确时间窗 的去重存储。留作独立一单,不塞进这次修复。 **怎么验证它真的好了。** 在一个已经有引擎回合的会话里发 `/help`,把标签页 放着:十秒后输出还在,刷新之后还在。对着一个「回合正在跑」的服务器制造一次 -确认超时:错误条会说引擎正在执行这条消息,且输入框是空的。 +确认超时:错误条会说引擎正在执行这条消息,且输入框是空的。等这个回合跑完: +灰色错误条自己消失。 + +## 输入区的状态按会话隔离(webui-parity 106) + +用户停在输入区的一切——正在打的文字、附件 chips、发送失败红条——都存在 +当前会话的键下(`webapp/lib/composer-draft.ts`,以 `state.sessionId` 为键 +的 `Map`;`""` 是首页无会话的桶)。切换会话就是换一个盒子:会话 B 永远 +不会显示会话 A 的草稿或红条,来回切换两边的状态都不丢。质检 s28 截图 +拍到的正是这个存储的共享桶版本:会话 2 的视图同时挂着会话 1 的草稿、 +409 红条和模型 chip。 + +按会话存储、而不是「切换时清空」,是权衡后的决定:清空 effect(#141 的 +过渡修法)在用户**切回来**时同样触发,恰恰毁掉他们回来要看的那份草稿和 +没读完的红条。按会话分键保住了旧全局行为里好的那一半(来回跳会话什么都不 +丢),又去掉了串扰。草稿不落 `localStorage`——它们是本次页面访问的工作 +状态;持久化面仍归 `lib/persist.ts` 的契约管。 + +模型选择器 chip 的**值**一直读服务端快照,本就不需要隔离;它的本地 UI +状态(打开的级联、预览中的行、按模型记的草稿镜像)在会话键变化时重置, +会话 A 的菜单状态不会在会话 B 的视图里残留。至于在一个会话视图里做的 +模型选择会不会落进另一个会话的引擎配置,那是服务端 +`applyConfigOptionUpdate` 的事,不在本单前端范围内。 + +**怎么验证它真的好了。** 在会话 A 打一段草稿,切到会话 B:B 的输入框是 +空的,chip 跟着 B 的服务端模型走。切回 A:草稿和没读完的红条原样都在。 + ## 端点清单(依据当前源码) diff --git a/packages/webui/docs/ARCHITECTURE.md b/packages/webui/docs/ARCHITECTURE.md index 6a2638c4..b75babf0 100644 --- a/packages/webui/docs/ARCHITECTURE.md +++ b/packages/webui/docs/ARCHITECTURE.md @@ -489,15 +489,28 @@ not import it but adopts the same shape. Unknown future statuses render as ### `engine/` (capability declarations + the local-runtime-v2 host) The engine abstraction lives at `server/engine/` (engine-abstraction -batch B1; migration state M1). Five files, one job each: +batch B1; migration state M1, plus M3 batches B0, B1, B2 and B3). Eleven +files, one job each: | File | Owns | | --- | --- | | `engine/capabilities.js` | The contract: `ENGINE_CAPABILITY_KEYS` (the 14 matrix keys), `validateEngineCapabilities`, `assertEngineCapability`, `summarizeUnavailableCapabilities` | | `engine/errors.js` | `EngineCapabilityNotSupportedError` + `engineCapabilityHttpResponse` (the 501 payload shape) | -| `engine/index.js` | The facade: `getEngineProvider`, `listEngineProviderIds` (registry by provider id; transport selection arrives with migration step M4) | -| `engine/providers/local-runtime-v2.js` | `createCatalogueHost` (moved verbatim from `runtime-host.js`, which re-exports it) + `LOCAL_RUNTIME_V2_CAPABILITIES` | +| `engine/host.js` | `getEngineCatalogueHost` — the lazy bridge to the one catalogue host. No static import of the host module: the getter body is a dynamic `import()` of `lib/acp-client.js`, so the facade costs a function, not a module load | +| `engine/index.js` | The facade: `getEngineProvider`, `listEngineProviderIds`, `getEngineCatalogueHost` (registry by provider id; transport selection arrives with migration step M4) | +| `engine/providers/local-runtime-v2.capabilities.js` | `LOCAL_RUNTIME_V2_CAPABILITIES` — **declaration only, and the split is load-bearing**: its sole import is `../capabilities.js`, so `/api/engine-capabilities` can read the capability table without pulling the v2 host's TypeScript dependency tree (~4.7 s of first-compile) into the boot path. That tree stays behind the same lazy boundary `acp-client.js` already documented | +| `engine/providers/local-runtime-v2.js` | `createCatalogueHost` (moved verbatim from `runtime-host.js`, which re-exports it) + re-exports the declaration above, so consumers keep one import shape. This is the heavy one — `@mavis/local-runtime-v2`, `@mavis/config`, `@minimax/code/runtime-adapter` — and no file `app.js` reaches may import it | | `engine/providers/tui-runtime-adapter.js` | `TUI_RUNTIME_ADAPTER_CAPABILITIES` (declaration only — the adapter itself is constructed inside the v2 host) | +| `engine/session-reads.js` | The directory-read family's facade calls (`readEngineSessionList`, `readEngineSessionListForWorkspace`, `readEngineSessionTitle`, `readEngineVersion`) and the endpoint→capability table `SESSION_READ_ENDPOINTS` (step M3, batch B1) | +| `engine/session-tree-reads.js` | The session-tree family's facade call (`readEngineSessionTree`) and the endpoint→capability table `SESSION_TREE_ENDPOINTS` (step M3, batch B2). Gates **hard**: `assertSessionTreeCapability` throws → 501, because the tree is entirely engine data. Forwards to `lib/session-tree.js#getSessionTree`; the assembler is not duplicated | +| `engine/session-export.js` | The export family's facade call (`readEngineSessionTranscript`) and the endpoint→capability table `SESSION_EXPORT_ENDPOINTS` (step M3, batch B2). Gates **soft**: `checkSessionExportCapability` reports and never throws, because export's primary source is `sessions.json`, not the engine | +| `engine/usage-reads.js` | The usage family's facade calls (`readEngineAccountQuota`, `readEngineSessionUsage`, `readEngineQuotaForecast`), the derived figure `contextUsedTokens`, and the endpoint→capability table `USAGE_READ_ENDPOINTS` (step M3, batch B3). Gates **hard** on the two engine reads and declares **no capability at all** for #19, which touches no engine surface | + +Routes take the host from the facade and never from `lib/acp-client.js`: +`routes/plugins.js` and `routes/turn-diff.js` call +`getEngineCatalogueHost()`. Both keep a `deps`-injected data source +(`deps.getCliService`, `deps.getDiffApplication`) so the handler suites stay +hermetic. Declaration discipline (admission rules for any future provider, enforced by the snapshot tests in `test/lib/engine/capabilities.test.js`): @@ -515,17 +528,227 @@ by the snapshot tests in `test/lib/engine/capabilities.test.js`): forbidden** — a missing capability must be legible before the call and loud after it (#110 fake-success discipline). 4. One host per provider process-wide: `createCatalogueHost` remains the - single owner of the runtime instance (`acp-client.js#getCatalogueHost` - keeps its "Never build a second host" rule); `close()` stays bounded. + single owner of the runtime instance, and the only way to reach it is the + facade's `getEngineCatalogueHost()` (which forwards to + `acp-client.js#getCatalogueHost` and its "Never build a second host" rule); + `close()` stays bounded. Two `CliService` instances over one dataDir is a + split brain against the plugin / local-disable tables, not a redundancy. 5. Levels drive the UI, never provider names: the frontend reads `GET /api/engine-capabilities` (`routes/engine-capabilities.js#handleEngineCapabilities`) and renders `full` / `partial`(+missing) / `none` — no hard-coded provider lists in UI code. +### Declaration-vs-implementation snapshot (M2) + +A declaration is only as honest as the check behind it. +`test/lib/engine/capability-snapshot.test.js#auditProviderCapabilities` +audits every `full`/`partial` key of both registered providers against a +REAL catalogue host booted once per run on an isolated tmp data dir +(`MINIMAX_DATA_DIR` plus every `MCODE_WEBUI_*` path pinned BEFORE the +provider import — setting only `MCODE_WEBUI_DATA_DIR` would leave the +engine dir falling back to `~/.minimax` and rewriting the user's real +config): + +- `full` — every tracked method of the key must be a function on the + declared surface member (`adapter`, `cliService`, or + `applications.session.diff`); +- `partial` — the present half must exist; every method-named `missing` + item must be genuinely absent; an absent method that dropped out of + `missing` goes red (under-declaration); and kebab-case sub-capability + names (`file-write`, `git-diff`, …) go red the moment a covering + method appears on the surface — a future `getWorkspaceGitDiff` forces + the `git-diff` entry to be re-audited; +- `none` — deliberately not method-checked; a provider may expose no + surface for the capability. + +The tracked method table (`REQUIRED_METHODS` in the same file) was +derived from the live surfaces themselves (prototype-chain reflection: +91 adapter methods, 94 CliService methods, the session.diff facade), not +copied from the design matrix. The audit is a pure function over +(declaration, method sets), and the mutation tests in the same file pin +that each drift class — a flipped level, a deleted method, a grown +sub-capability — turns it red. A registry-driven static guard sweeps +every REGISTERED provider (`engine/index.js#listEngineProviderIds`) for +the exact 14-key set, so a typo'd or unknown key cannot pass silently, +and providers registered by M4 will be swept without editing the test. + Runtime probing (downgrading a declared level when the environment disagrees) is deliberately absent in this batch — see `engine/index.js` for the reasoning. +Boot-path discipline: `app.js` reaches `engine/index.js`, so that file and +everything it imports statically must stay free of `@mavis/*`, +`@minimax/*` and the host modules. M1 learned that by paying for it +(209ms → 2700ms at server start; the facade's own load 4685ms → 5ms after +declaration and construction were split). `test/lib/engine/host-facade.test.js` +enforces it against the real module graph rather than against source text. +`engine/session-reads.js`, `engine/session-tree-reads.js`, +`engine/session-export.js` and `engine/usage-reads.js` all live under the +same rule: their static imports are `engine/capabilities.js` and +`engine/index.js` only, and every heavier dependency — +`lib/acp-client.js`, `lib/config.js`, `lib/session-tree.js`, +`lib/transcript.js`, `lib/usage.js`, `lib/mavis-usage.js` and +`lib/quota-forecast.js` — is reached through `await import()` inside the +functions. + +#### Which endpoints read through the facade (step M3, batch B1) + +`engine/session-reads.js` covers the five directory-read endpoints. Each +row names the capability it gates on and the provider method it depends +on, so a `partial` declaration that drops exactly that method answers 501 +naming it: + +| Endpoint | Capability · sub-item | Value source | +| --- | --- | --- | +| `GET /api/acp-sessions` | `sessionCrud` · `listSessions` | `acp-client.js#getMcodeSessionsForWorkspace` (30s cache, cwd normalisation) | +| `GET /api/acp-session-title` | `sessionCrud` · `getSession` | `acp-client.js#getMcodeSessionTitle` | +| `GET /api/protocol/list-sessions` | `sessionCrud` · `listSessions` | `acp-client.js#listAllMcodeSessions`; the route keeps its own cwd filter | +| `GET /api/state` | `sessionCrud` · `listSessions` | the `mcodeSessions` mirror only — `snapshotViewFields` / `mcodeSessionsSnapshotFields` are untouched | +| `GET /api/health` | none of the 14 keys | the ACP `initialize` `agentInfo.version` mirror; the catalogue host exposes no version accessor, so the facade reports the source instead of inventing one | + +Three properties this layer holds, each with a test behind it: + +1. **One normalizer.** The runtime path is projected by + `lib/catalogue-sessions.js#projectTuiSessionToAcp`, which mirrors the + ACP adapter's `toAcpSessionInfo` rule for rule — `title` and + `updatedAt` are omitted when absent, never emitted as `null`. The + facade forwards that projection; it does not re-project it. +2. **Where the bytes came from is reported, not assumed.** Every read + answers a `source` of `catalogue`, `acp` or `acp-fallback` (the + transport asked for the catalogue host and got `null`). It is + metadata, not wire — the endpoints' payloads are byte-identical before + and after the facade. +3. **The gate is real.** The registered provider declares `sessionCrud` + `full`, so nothing 501s today; the tests drive a fixture declaration + that lacks `listSessions` and assert the 501 payload. A gate nobody + ever exercises is indistinguishable from no gate. + +#### Which endpoints read through the facade (step M3, batch B3) + +`engine/usage-reads.js` covers the four usage endpoints (#15, #16, #17, +#19). This family is where a refactor can be entirely silent, because three +of its four numbers are derived rather than counted — so the table below is +as much about where each number comes from as about which capability gates +it: + +| Endpoint | Capability · sub-item | Value source | +| --- | --- | --- | +| `POST /api/usage` | `authCredentials` · `getAccountStatus` | `lib/usage.js#runUsageQuery` — the engine's `mcode/account/status` projection, copied into `cs.usage`; the payload is written byte-for-byte, `ok:false` / `error` shape included | +| `POST /api/usage-trigger` | `authCredentials` · `getAccountStatus` | the same read; the two endpoints differ only in the client's `record` flag, which is the difference between a reading and a measurement | +| `GET /api/usage-real` | `usageStats` · `getSessionUsage` | `lib/mavis-usage.js` over the engine's own `local_runtime_token_usage` table. `contextUsed` is derived here by `contextUsedTokens` | +| `GET /api/usage/forecast` | none of the 14 keys | webui's own `~/.mcode-webui/usage-history.ndjson`, via `lib/quota-forecast.js`. It calls no engine surface, so it declares none | + +Four properties this family holds, each with a test behind it: + +1. **`contextUsed` is cumulative, and the cache counters are not in it.** + `totalInput + totalOutput + totalReasoning`. The cache counters are a + SUBSET of `input`, so adding them double-counts; `totalCacheWrite` is + not part of the context window at all. This is also NOT the chat flow's + `lastTurnContextTokens`: the context bar shows one turn's worth, `#17` + shows the session's spend, and `test/lib/engine/usage-reads.test.js` + perturbs each of the seven numeric fields one at a time so a merged or + "simplified" formula flips a row instead of quietly shipping. +2. **`totalReasoning` is the database's own `SUM`, forwarded.** The + snapshot test reads the same aggregate with plain SQL and compares; a + facade that re-derived it from anything else fails. +3. **The forecast is a pure function of a history prefix.** Every prefix of + a growing history is compared against the module's own + `forecastExhaustion(readHistory())` at the same instant, and the sample + count's flat stretch across the deliberately-null sample is asserted, so + a read that re-filtered, re-sorted or re-sampled would break the + sequence rather than the shape. +4. **A `none` / `partial`-missing declaration would 501.** The registered + provider declares both `authCredentials` and `usageStats` `full`, so only + the fixture-driven tests can prove the gate bites. #19's `null` row is + the counter-example with a reason: gating a read that touches no engine + surface would remove a working endpoint in response to a declaration + about something it does not depend on. + +`#17` declares `usageStats` · `getSessionUsage` but does not yet CALL that +method; it reads the same SQLite table the method reads, through +`lib/mavis-usage.js`. Three measured reasons, stated in the module header: +the catalogue host only exists under the `runtime` transport +(`acp-client.js#transportWantsCatalogue`), and `acp` is the default; +`getSessionUsage` answers `{summary, rows: UsageView[]}` where the endpoint +answers a per-column aggregate with `rows` as a COUNT, so switching would +mean rebuilding `totalReasoning` and `contextUsed` from a different +starting point; and it would put the v2 TypeScript tree on an endpoint that +needs nothing from it. M4 is where the two are allowed to meet. + +The transport→provider table has one entry (`runtime`). Under the default +`acp` transport no provider is registered yet, so the gate reports +`unregistered-transport` and passes through — M4 registers the ACP +provider and the table gains its row. Passing through is not the same as +claiming support, and the two are reported differently on purpose. + +#### Which endpoints read through the facade (step M3, batch B2) + +Batch B2 adds two endpoints, and they are the first two whose gate policies +**differ**. They are separate files for that reason; merging them would force +one to inherit the other's. + +| Endpoint | Capability · sub-item | Enforcement | Value source | +| --- | --- | --- | --- | +| `GET /api/session-tree` | `sessionCrud` · `listSessions` | hard — 501 | `lib/session-tree.js#getSessionTree`, forwarded verbatim | +| `GET /api/sessions/:id/export` | `sessionCrud` · `getSession` | soft — reported | `lib/transcript.js#readMcodeTranscript` (the enrichment only) | + +**Why one gate throws and the other does not.** `/api/session-tree` is +entirely engine data: the hierarchy is assembled from `local_runtime_sessions` +in the runtime db, so a provider that cannot list sessions genuinely has no +tree to return, and 501 is the honest answer. `/api/sessions/:id/export` is +mostly *not* engine data — the conversation comes from `sessions.json`, and +the engine only contributes a best-effort transcript enrichment the endpoint +has always promised never to block on. Gating it hard would delete working +functionality in response to a declaration about a capability the endpoint +does not depend on. So `checkSessionExportCapability` answers what the +provider declared and returns; the caller degrades `_meta.mcode_unavailable` +through the endpoint's own pre-existing channel, and the export still serves +the full webui chat. `test/lib/engine/session-export.test.js` pins this by +swapping in a provider that declares `sessionCrud: none` and asserting that +export reports while the tree family throws on the same fixture. + +Four properties this batch holds, each with a test behind it: + +1. **The node shape is unchanged, and it is asymmetric.** A root node + carries `{id, title, agent, kind, status, updatedAt, children}`; a child + node carries the same fields **without** `children`, because + `buildTree` adds that key only in the output map that wraps each root. + Measured on the real tree: 233 root nodes carry `children`, all 66 child + nodes do not. "Normalising" this would change 66 nodes' shape in the + sidebar. +2. **There is no `parent_session_id` in the response.** The hierarchy is + structural — expressed through `children` — and `parent_session_id` + exists only inside the db read. A future addition of that key to the node + is a client-visible change, so the exact key set is asserted per depth. +3. **One assembler.** `buildTree` remains the only thing that decides which + rows attach to which parent, and the route does not re-derive the + hierarchy. Rows that cannot attach — an orphan whose parent is not in the + row set, a cross-directory parent, a grandchild, a child of a `root` + container row, anything in a cycle — are dropped, as they always were. + That is why the batch was verified by exporting the tree before and + after and diffing every node, not by counting rows. +4. **The tree's 501 is not swallowed.** The route's existing `try/catch` + would otherwise fold the capability error into its own + `{ok:false, reason:"session_tree_failed"}` soft-fail body and turn a 501 + into a 200. The route re-throws `EngineCapabilityNotSupportedError` and + keeps the soft-fail path for everything else. + +**`source` is not transport-switched for the tree.** The tree is read from +the engine's own runtime db, which both the `runtime` and the `acp` +transport can see, so `readEngineSessionTree` reports `source: "runtime-db"` +under every transport rather than claiming a catalogue answer. The +declaration check is still transport-keyed: which provider is active is a +transport question even when the read itself is not. + +**Export's enrichment is currently inert against the v2 schema, by +design.** `lib/transcript.js` keeps its `v2-data-json` probe OUT of the +default probe set so that export's behaviour does not change, and the live +`local_runtime_message_rows` has no `content` column. So on a current +runtime db the enrichment answers `no_matching_table` and every export +reports `_meta.mcode_unavailable: true` with +`_meta.source: "webui"`. That is pre-existing and deliberately preserved — +re-enabling it is a behaviour change for a later slice, not a refactor. + ## 4. The `clientState` payload This is the shape every SSE `state` event contains. The webui mirrors diff --git a/packages/webui/docs/ARCHITECTURE.zh-CN.md b/packages/webui/docs/ARCHITECTURE.zh-CN.md index 324ea6a7..3bce7f75 100644 --- a/packages/webui/docs/ARCHITECTURE.zh-CN.md +++ b/packages/webui/docs/ARCHITECTURE.zh-CN.md @@ -461,15 +461,26 @@ queued \| done \| stopped`)是投影层产物、不是存储值;webui 不导 ### `engine/`(能力声明 + local-runtime-v2 host) 引擎抽象层位于 `server/engine/`(engine-abstraction 批次 B1;迁移 -状态 M1)。五个文件,各管一件事: +状态 M1,外加 M3 的 B0、B1、B2 与 B3 四批)。十一个文件,各管一件事: | 文件 | 职责 | | --- | --- | | `engine/capabilities.js` | 契约本体:`ENGINE_CAPABILITY_KEYS`(14 个矩阵键)、`validateEngineCapabilities`、`assertEngineCapability`、`summarizeUnavailableCapabilities` | | `engine/errors.js` | `EngineCapabilityNotSupportedError` 与 `engineCapabilityHttpResponse`(501 载荷形状) | -| `engine/index.js` | 门面:`getEngineProvider`、`listEngineProviderIds`(按 provider id 的注册表;按 `MCODE_WEBUI_TRANSPORT` 选传输在迁移步 M4 引入) | -| `engine/providers/local-runtime-v2.js` | `createCatalogueHost`(自 `runtime-host.js` 原样移入,后者转发导出)+ `LOCAL_RUNTIME_V2_CAPABILITIES` | +| `engine/host.js` | `getEngineCatalogueHost`——通往那唯一 catalogue host 的惰性桥。对 host 模块零静态 import:函数体里是 `lib/acp-client.js` 的动态 `import()`,所以门面付出的是一个函数,不是一次模块加载 | +| `engine/index.js` | 门面:`getEngineProvider`、`listEngineProviderIds`、`getEngineCatalogueHost`(按 provider id 的注册表;按 `MCODE_WEBUI_TRANSPORT` 选传输在迁移步 M4 引入) | +| `engine/providers/local-runtime-v2.capabilities.js` | `LOCAL_RUNTIME_V2_CAPABILITIES`——**只有声明,且这个拆分是有承重意义的**:它唯一的 import 是 `../capabilities.js`,所以 `/api/engine-capabilities` 读能力表时**不会把 v2 host 的 TypeScript 依赖树(首次编译约 4.7 秒)拖进 boot 路径**。那棵依赖树仍留在 `acp-client.js` 早已注明的 lazy 边界之后 | +| `engine/providers/local-runtime-v2.js` | `createCatalogueHost`(自 `runtime-host.js` 原样移入,后者转发导出)+ 转发导出上面的声明,消费方的 import 形状因此不变。它是重的那一个——`@mavis/local-runtime-v2`、`@mavis/config`、`@minimax/code/runtime-adapter`——`app.js` 能触达的文件里绝不许 import 它 | | `engine/providers/tui-runtime-adapter.js` | `TUI_RUNTIME_ADAPTER_CAPABILITIES`(仅声明——adapter 本体在 v2 host 内构造) | +| `engine/session-reads.js` | 目录读族的面板调用(`readEngineSessionList`、`readEngineSessionListForWorkspace`、`readEngineSessionTitle`、`readEngineVersion`)与端点→能力对照表 `SESSION_READ_ENDPOINTS`(迁移步 M3 批次 B1) | +| `engine/session-tree-reads.js` | 会话树族的面板调用 `readEngineSessionTree` 与端点→能力对照表 `SESSION_TREE_ENDPOINTS`(迁移步 M3 批次 B2)。**硬门控**:`assertSessionTreeCapability` 抛出 → 501,因为树完全由引擎数据构成。转发到 `lib/session-tree.js#getSessionTree`,树的装配逻辑不复制第二份 | +| `engine/session-export.js` | 导出族的面板调用 `readEngineSessionTranscript` 与端点→能力对照表 `SESSION_EXPORT_ENDPOINTS`(迁移步 M3 批次 B2)。**软门控**:`checkSessionExportCapability` 只报告、从不抛出,因为导出的主数据源是 `sessions.json` 而非引擎 | +| `engine/usage-reads.js` | 用量族的面板调用(`readEngineAccountQuota`、`readEngineSessionUsage`、`readEngineQuotaForecast`)、派生量 `contextUsedTokens`,与端点→能力对照表 `USAGE_READ_ENDPOINTS`(迁移步 M3 批次 B3)。两个引擎读**硬门控**;#19 **完全不声明能力**,因为它不触达任何引擎面 | + +路由从门面取 host,不从 `lib/acp-client.js` 取:`routes/plugins.js` 与 +`routes/turn-diff.js` 调 `getEngineCatalogueHost()`。两者都保留 `deps` +注入的数据源(`deps.getCliService`、`deps.getDiffApplication`), +handler 层测试因此保持封闭。 声明纪律(未来任何 provider 的准入规则,由 `test/lib/engine/capabilities.test.js` 的快照测试强制): @@ -484,17 +495,189 @@ queued \| done \| stopped`)是投影层产物、不是存储值;webui 不导 `501 engine_capability_not_supported`。**禁止空实现**——缺能力必须在 调用前可读、调用后响亮(#110 假成功纪律)。 4. 每 provider 进程内单 host:`createCatalogueHost` 仍是运行时实例的 - 唯一所有者(`acp-client.js#getCatalogueHost` 的「绝不建第二个 host」 - 规则不变);`close()` 保持有界。 + 唯一所有者,触达它的唯一入口是门面的 `getEngineCatalogueHost()` + (转发到 `acp-client.js#getCatalogueHost`,其「绝不建第二个 host」 + 规则不变);`close()` 保持有界。同一 dataDir 上两个 `CliService` + 实例是对 plugin / local-disable 表的脑裂,不是冗余。 5. 驱动 UI 的是档位,不是 provider 名单:前端读 `GET /api/engine-capabilities` (`routes/engine-capabilities.js#handleEngineCapabilities`), 按 `full` / `partial`(+missing)/ `none` 三档渲染——UI 代码里不出现 硬编码的 provider 名单。 +### 声明与实现的快照校验(M2) + +声明有多诚实,取决于背后的校验有多硬。 +`test/lib/engine/capability-snapshot.test.js#auditProviderCapabilities` +对两个已注册 provider 的每个 `full`/`partial` 键做审计,对象是**真实** +的 catalogue host——每次运行在隔离的临时数据目录上起一个 +(`MINIMAX_DATA_DIR` 与全部 `MCODE_WEBUI_*` 路径在 provider import +**之前**钉死;只设 `MCODE_WEBUI_DATA_DIR` 不够,引擎目录会回落到 +`~/.minimax` 改写用户真实配置): + +- `full`——该键跟踪的方法必须在声明的 surface 成员上 + (`adapter`、`cliService` 或 `applications.session.diff`)全部为函数; +- `partial`——存在的部分必须在;方法名形态的 `missing` 项必须真的 + 不存在;某缺席方法从 `missing` 里被拿掉会红(声明不完整);kebab-case + 子能力名(`file-write`、`git-diff` 等)在 surface 上出现覆盖方法的那一刻 + 变红——将来引擎长出 `getWorkspaceGitDiff`,`git-diff` 这条就必须重新审计; +- `none`——刻意不做方法校验;provider 允许对该能力完全不设接口面。 + +方法跟踪表(同文件内的 `REQUIRED_METHODS`)取自真实 surface 本身 +(原型链反射:adapter 91 个方法、CliService 94 个、session.diff 门面), +不是从设计矩阵抄的。审计是对(声明, 方法集)的纯函数,同文件的变异测试 +钉住每类漂移——改档位、删方法、子能力长出方法——各自必然变红。另有 +注册表驱动的静态守卫扫过每个**已注册** provider +(`engine/index.js#listEngineProviderIds`)的 14 键集合,拼错或多写的键 +无法静默通过;M4 注册 acp/exec provider 时无需改测试即被覆盖。 + 运行时探测(环境不符时把声明档位降级)本批刻意未做——理由见 `engine/index.js` 头注释。 +启动路径纪律:`app.js` 会触达 `engine/index.js`,因此该文件及其全部 +静态依赖必须不含 `@mavis/*`、`@minimax/*` 与任何 host 模块。M1 是交过 +学费才换来这条(server 启动 209ms → 2700ms;声明与构造拆成两个文件后, +门面自身加载 4685ms → 5ms)。`test/lib/engine/host-facade.test.js` +对着真实模块图强制它,而不是对着源码文本。 +`engine/session-reads.js`、`engine/session-tree-reads.js`、 +`engine/session-export.js` 与 `engine/usage-reads.js` 全部服从同一条 +纪律:静态 import 只有 `engine/capabilities.js` 与 `engine/index.js`, +而每个更重的依赖——`lib/acp-client.js`、`lib/config.js`、 +`lib/session-tree.js`、`lib/transcript.js`、`lib/usage.js`、 +`lib/mavis-usage.js` 与 `lib/quota-forecast.js`——都在函数体内用 +`await import()` 触达。 + +#### 哪些端点走门面读(迁移步 M3 批次 B1) + +`engine/session-reads.js` 覆盖 5 个目录读端点。每一行写明它门控的 +能力键与它依赖的 provider 方法,因此一份恰好缺该方法的 `partial` +声明会 501 并点名是哪个方法: + +| 端点 | 能力 · 子项 | 取值来源 | +| --- | --- | --- | +| `GET /api/acp-sessions` | `sessionCrud` · `listSessions` | `acp-client.js#getMcodeSessionsForWorkspace`(30s 缓存 + cwd 归一化) | +| `GET /api/acp-session-title` | `sessionCrud` · `getSession` | `acp-client.js#getMcodeSessionTitle` | +| `GET /api/protocol/list-sessions` | `sessionCrud` · `listSessions` | `acp-client.js#listAllMcodeSessions`;cwd 过滤仍留在路由里 | +| `GET /api/state` | `sessionCrud` · `listSessions` | 只作用于 `mcodeSessions` 镜像——`snapshotViewFields` / `mcodeSessionsSnapshotFields` 一字未动 | +| `GET /api/health` | 14 键中无对应键 | ACP `initialize` 的 `agentInfo.version` 镜像;catalogue host 没有版本访问器,面板如实报告来源而不是凭空造一个方法 | + +本层守住三条性质,每条背后都有测试: + +1. **只有一个 normalizer。** runtime 路径由 + `lib/catalogue-sessions.js#projectTuiSessionToAcp` 投影,逐条镜像 + ACP adapter 的 `toAcpSessionInfo` 规则——`title` 与 `updatedAt` + 缺失时**省略该键**,绝不输出 `null`。面板原样转发这份投影,不做 + 二次投影。 +2. **字节来自哪里是报告出来的,不是假设的。** 每次读都回答一个 + `source`:`catalogue` / `acp` / `acp-fallback`(传输要了 catalogue + host 但拿到 `null`)。它是元数据,不上线——端点载荷在接面板前后 + 逐字节相同。 +3. **门控是真的。** 已注册的 provider 声明 `sessionCrud` 为 `full`, + 所以今天没有任何端点会 501;测试用一份缺 `listSessions` 的样本声明 + 驱动出 501 载荷。没人跑过的门控与没有门控无法区分。 + +#### 哪些端点走门面读(迁移步 M3 批次 B3) + +`engine/usage-reads.js` 覆盖 4 个用量端点(#15、#16、#17、#19)。 +这一族是「重构全程静默」的重灾区:四个数字里有三个是**算出来的** +而不是数出来的,所以下表不只写门控哪个能力,更写清每个数字从哪来: + +| 端点 | 能力 · 子项 | 取值来源 | +| --- | --- | --- | +| `POST /api/usage` | `authCredentials` · `getAccountStatus` | `lib/usage.js#runUsageQuery`——引擎的 `mcode/account/status` 投影,抄进 `cs.usage`;载荷逐字节写出,含 `ok:false` / `error` 形状 | +| `POST /api/usage-trigger` | `authCredentials` · `getAccountStatus` | 同一次读;两个端点只差客户端的 `record` 标志,而它决定这次是「读数」还是「采样」 | +| `GET /api/usage-real` | `usageStats` · `getSessionUsage` | `lib/mavis-usage.js` 读引擎自己的 `local_runtime_token_usage` 表;`contextUsed` 由 `contextUsedTokens` 在此派生 | +| `GET /api/usage/forecast` | 14 键中无对应键 | webui 自己的 `~/.mcode-webui/usage-history.ndjson`,经 `lib/quota-forecast.js`。它不触达任何引擎面,所以不声明任何能力 | + +本层守住四条性质,每条背后都有测试: + +1. **`contextUsed` 是累计值,且不含缓存计数。** 公式是 + `totalInput + totalOutput + totalReasoning`。缓存计数是 `input` 的 + **子集**,加上会重复计数;`totalCacheWrite` 根本不在上下文窗口里。 + 它也**不是**聊天流程的 `lastTurnContextTokens`:上下文条显示的是 + 一轮的量,`#17` 显示的是整会话的花费。 + `test/lib/engine/usage-reads.test.js` 对七个数值字段逐个扰动, + 被合并或被「简化」的公式会翻掉某一行,而不是悄悄发版。 +2. **`totalReasoning` 是数据库自己的 `SUM`,原样转发。** 快照测试用 + 裸 SQL 独立算出同一个聚合再比对;门面若从别处重新派生,此测试即红。 +3. **预测是历史前缀的纯函数。** 增长中的历史的每一个前缀,都在同一时刻 + 与模块自己的 `forecastExhaustion(readHistory())` 比对,并且断言样本数 + 在那条故意置 `null` 的样本处出现的「平台期」——所以重新过滤、重新排序 + 或重新采样会破坏**序列**而不只是破坏形状。 +4. **`none` / 缺子项的 `partial` 声明会 501。** 已注册的 provider 把 + `authCredentials` 与 `usageStats` 都声明为 `full`,所以只有样本驱动 + 的测试能证明门控会咬。#19 那一行 `null` 是带理由的反例:给一个 + 根本不触达引擎面的读加硬门控,等于用一条与它无关的声明去关掉一个 + 正常工作的端点。 + +`#17` 声明了 `usageStats` · `getSessionUsage`,但**尚未调用**该方法: +它经 `lib/mavis-usage.js` 读的是该方法读的同一张 SQLite 表。三条实测 +理由写在模块头注释里——catalogue host 只在 `runtime` 传输下存在 +(`acp-client.js#transportWantsCatalogue`),而 `acp` 是默认值; +`getSessionUsage` 回答的是 `{summary, rows: UsageView[]}`,端点回答的是 +按列聚合且 `rows` 是 COUNT 的形状,换过去就意味着从另一个起点重建 +`totalReasoning` 与 `contextUsed`;而且它会把 v2 的 TypeScript 依赖树压到 +一个本来不需要它的端点的应答路径上。M4 才是两者允许会合的地方。 + +传输→provider 表目前只有 `runtime` 一条。默认 `acp` 传输下尚无已注册 +provider,于是门控报告 `unregistered-transport` 并放行——M4 注册 ACP +provider 后该表补上对应行。放行不等于声称支持,二者刻意分开报告。 + +#### 哪些端点走门面读(迁移步 M3 批次 B2) + +批次 B2 收编 2 个端点,它们是前两个**门控策略不同**的端点。正因如此才 +拆成两个文件:合并会迫使其中一个继承另一个的策略。 + +| 端点 | 能力 · 子项 | 强制方式 | 取值来源 | +| --- | --- | --- | --- | +| `GET /api/session-tree` | `sessionCrud` · `listSessions` | 硬——501 | `lib/session-tree.js#getSessionTree`,原样转发 | +| `GET /api/sessions/:id/export` | `sessionCrud` · `getSession` | 软——只报告 | `lib/transcript.js#readMcodeTranscript`(仅增强部分) | + +**为什么一个门控抛错、另一个不抛。** `/api/session-tree` 完全是引擎数据: +层级由运行时库 `local_runtime_sessions` 装配,所以一个列不出会话的 +provider 确实没有树可返回,501 才是诚实答案。 +`/api/sessions/:id/export` 则**主要不是**引擎数据——对话来自 +`sessions.json`,引擎只贡献一份尽力而为的 transcript 增强,而该端点一直 +承诺绝不因此阻断导出。把它改成硬门控,等于因为一条关于「本端点并不依赖的 +能力」的声明而删掉本来能用的功能。所以 `checkSessionExportCapability` +只回答 provider 声明了什么然后返回;调用方通过端点既有的通道降级 +`_meta.mcode_unavailable`,导出照旧完整返回 webui 的对话。 +`test/lib/engine/session-export.test.js` 用一份声明 `sessionCrud: none` +的 provider 钉住这一点:同一份样本下,导出族报告、树族抛错。 + +本批守住的四条性质,每条背后都有测试: + +1. **节点形状未变,而且它是不对称的。** 根节点带 + `{id, title, agent, kind, status, updatedAt, children}`;子节点带同样 + 这些字段但**没有** `children`——因为 `buildTree` 只在包裹每个根节点的 + 输出映射里补这个键。在真实树上实测:233 个根节点带 `children`, + 66 个子节点全都不带。「顺手规范化」会让侧边栏里 66 个节点的形状改变。 +2. **响应里没有 `parent_session_id`。** 层级是结构性的——由 `children` + 表达——`parent_session_id` 只存在于读库阶段。将来把这个键加到节点上 + 就是客户端可见的变更,所以测试按深度逐字断言键集合。 +3. **只有一个装配器。** `buildTree` 仍是唯一决定哪些行挂到哪个父节点 + 的地方,路由不重新推导层级。挂不上的行——父节点不在结果集里的孤儿、 + 跨目录的父节点、孙节点、挂在 `root` 容器行下的子节点、任何处于环中的 + 行——照旧被丢弃。正因如此,本批的验证方式是改前改后各导一次树、 + 逐节点比对,而不是数行数。 +4. **树的 501 不被吞掉。** 路由原有的 `try/catch` 否则会把能力错误 + 折进它自己的 `{ok:false, reason:"session_tree_failed"}` 软失败体里, + 把 501 变成 200。路由重新抛出 `EngineCapabilityNotSupportedError`, + 其余错误仍走软失败。 + +**树的 `source` 不随传输切换。** 树读自引擎自己的运行时库,`runtime` +与 `acp` 两种传输都看得到,所以 `readEngineSessionTree` 在任何传输下都 +报告 `source: "runtime-db"`,而不是假称拿到了目录宿主。声明检查仍按传输 +分派:当前哪个 provider 生效是传输问题,即使这次读本身不是。 + +**export 的增强在 v2 表结构下当前是失效的,且这是刻意为之。** +`lib/transcript.js` 把 `v2-data-json` 探针留在默认探针集**之外**, +以保证 export 的行为不变;而线上真实的 `local_runtime_message_rows` +根本没有 `content` 列。因此在当前运行时库上增强会返回 +`no_matching_table`,每次导出都报告 `_meta.mcode_unavailable: true` 与 +`_meta.source: "webui"`。这是既有行为且被刻意保留——重新启用它是一次行为 +变更,属于后续切片,不属于这次收编。 + ## 4. `clientState` 载荷 这是每个 SSE `state` 事件所包含的形状。webui 将其 diff --git a/packages/webui/server/engine/host.js b/packages/webui/server/engine/host.js new file mode 100644 index 00000000..3050d1e4 --- /dev/null +++ b/packages/webui/server/engine/host.js @@ -0,0 +1,40 @@ +// webui/server/engine/host.js +// +// The lazy half of the engine facade: the one place route code asks for +// the live catalogue host (migration step M3, batch B0). +// +// Why the getter cannot simply live in `lib/acp-client.js` and be +// imported from `engine/index.js`: `getCatalogueHost()` is already lazy +// *inside* — it `await import("./runtime-host.js")` on first call — but +// the MODULE is not. `lib/acp-client.js` statically imports +// `../../acp.mjs`, the command registry, the settings/config chain and the +// session-delete module. `engine/index.js` is loaded by `app.js` at boot, +// so a static import of `acp-client.js` there would put the ACP client and +// everything behind it on every server start. That is the exact regression +// M1 already paid for once (209ms → 2700ms; index load 4685ms → 5ms after +// the declaration/construction split). This file exists to keep that +// boundary: the only thing `engine/index.js` gains is a function, and the +// function does not touch the module graph until it is called. +// +// Discipline, unchanged by the indirection: one host per process. This +// function FORWARDS to `acp-client.js#getCatalogueHost`, it does not +// construct anything. Two callers must never end up with two CliService +// instances over one dataDir — that is both wasteful and a split brain +// against the plugin / local-disable tables. +// +// The return value is passed through untouched, `null` included: a host +// that failed to boot is an answer (routes answer `RUNTIME_UNAVAILABLE`), +// never a licence to build a second one or to fall back to another path. + +/** + * The process-lifetime catalogue host, booted on first call. + * + * @returns {Promise} The host (the same object + * `lib/acp-client.js#getCatalogueHost` returns), or `null` when the + * runtime failed to boot. Errors thrown by the getter propagate + * unchanged — callers own the failure mapping. + */ +export async function getEngineCatalogueHost() { + const { getCatalogueHost } = await import("../lib/acp-client.js"); + return getCatalogueHost(); +} diff --git a/packages/webui/server/engine/index.js b/packages/webui/server/engine/index.js index 2237a5c8..ed0f6680 100644 --- a/packages/webui/server/engine/index.js +++ b/packages/webui/server/engine/index.js @@ -32,8 +32,12 @@ // // Migration state (design §2.4): M1 done — the host construction moved // into providers/local-runtime-v2.js and runtime-host.js re-exports it; -// no route's behaviour changed. M2–M4 will route new consumers through -// this facade one endpoint family at a time. +// no route's behaviour changed. M3's first batch (B0) done — the +// catalogue host itself is now reached through this facade too +// (engine/host.js), so the plugins and turn-diff routes no longer name +// lib/acp-client.js. M3 batches B1 (#9 #10 #72 #74 #75) and B3 (#15 #16 +// #17 #19) done. B2 (#8 #11) and the rest of M3, then M4, will route +// their consumers through this facade one endpoint family at a time. import { ENGINE_CAPABILITY_KEYS } from "./capabilities.js"; // Declarations only — importing the provider *host-construction* modules @@ -42,6 +46,9 @@ import { ENGINE_CAPABILITY_KEYS } from "./capabilities.js"; // Host construction stays behind the lazy boundary runtime-host.js // always had; nothing on the boot path may import // providers/local-runtime-v2.js or providers/acp.js-style host modules. +// The same rule applies one level up: engine/host.js reaches +// lib/acp-client.js through a dynamic import, so re-exporting it here +// costs a function, not a module load. import { LOCAL_RUNTIME_V2_CAPABILITIES } from "./providers/local-runtime-v2.capabilities.js"; import { TUI_RUNTIME_ADAPTER_CAPABILITIES } from "./providers/tui-runtime-adapter.js"; @@ -52,6 +59,67 @@ export { engineCapabilityHttpResponse, isEngineCapabilityNotSupportedError, } from "./errors.js"; +// The lazy host getter: a function definition, no host, no @mavis/* import. +export { getEngineCatalogueHost } from "./host.js"; +// The directory-read family's gated reads (step M3, batch B1). Re-exported +// here so the facade is the one import site for engine reads, but the +// dependency runs the other way too — session-reads.js consults +// getEngineProvider. That cycle is safe for one concrete reason: +// session-reads.js reads NOTHING from this module while it is being +// evaluated. Its own module-scope constant is a literal table, and every +// binding it needs from here (getEngineProvider, DEFAULT_ENGINE_PROVIDER_ID) +// is read inside a function body, so a cold `import("./engine/index.js")` +// can never hit a temporal dead zone. Keep it that way: a new top-level +// `const X = SOMETHING_FROM_INDEX` in session-reads.js breaks the re-export. +// It also stays off the boot path for the reason host.js does — +// lib/acp-client.js and lib/config.js are reached through dynamic import() +// inside the read functions. +export { + SESSION_READ_ENDPOINTS, + assertSessionReadCapability, + readEngineSessionList, + readEngineSessionListForWorkspace, + readEngineSessionTitle, + readEngineVersion, + resolveSessionReadProvider, +} from "./session-reads.js"; +// Step M3, batch B2: the session-tree read (#8) and the export +// enrichment read (#11). Two modules, not one, because their gate +// policies are opposite and a single file would force one of them to +// inherit the other's: #8 is 100% engine data and gates HARD (501 via +// `assertSessionTreeCapability`), while #11's primary source is +// `sessions.json` and gates SOFT (`checkSessionExportCapability` +// reports, never throws) so a provider that cannot serve a transcript +// degrades the enrichment instead of the export. The same TDZ rule as +// B1 applies to both: read nothing from this module at module scope. +export { + SESSION_TREE_ENDPOINTS, + assertSessionTreeCapability, + readEngineSessionTree, + resolveSessionTreeProvider, +} from "./session-tree-reads.js"; +export { + SESSION_EXPORT_ENDPOINTS, + checkSessionExportCapability, + readEngineSessionTranscript, + resolveSessionExportProvider, +} from "./session-export.js"; +// The usage family's gated reads (step M3, batch B3). Same cycle, same +// rule, same reasoning as session-reads.js above: usage-reads.js reads +// NOTHING from this module at module scope — its `USAGE_READ_ENDPOINTS` +// table is a literal and every binding it needs (`getEngineProvider`, +// `DEFAULT_ENGINE_PROVIDER_ID`) is read inside a function body. A new +// top-level `const X = SOMETHING_FROM_INDEX` in usage-reads.js breaks the +// re-export exactly as it would in session-reads.js. +export { + USAGE_READ_ENDPOINTS, + assertUsageReadCapability, + contextUsedTokens, + readEngineAccountQuota, + readEngineQuotaForecast, + readEngineSessionUsage, + resolveUsageReadProvider, +} from "./usage-reads.js"; export { LOCAL_RUNTIME_V2_CAPABILITIES } from "./providers/local-runtime-v2.capabilities.js"; export { TUI_RUNTIME_ADAPTER_CAPABILITIES } from "./providers/tui-runtime-adapter.js"; diff --git a/packages/webui/server/engine/session-export.js b/packages/webui/server/engine/session-export.js new file mode 100644 index 00000000..f02d44ec --- /dev/null +++ b/packages/webui/server/engine/session-export.js @@ -0,0 +1,235 @@ +// webui/server/engine/session-export.js +// +// Migration step M3, batch B2: the export enrichment read (导出增强读) — +// +// #11 GET /api/sessions/:id/export?format=md|json +// +// What this file is for. Export is the ONE endpoint in this migration +// whose primary data source is webui's own store, not the engine: the +// conversation comes from `lib/sessions.js` (`sessions.json`), and the +// route parses it with its own line grammar and renders md/json. The +// engine's only contribution is the transcript enrichment +// (`readMcodeTranscript`), which the route has always treated as +// BEST-EFFORT — "If the db is unreadable / the table is missing / the +// schema differs, we set `_meta.mcode_unavailable` and continue with the +// webui source — never block export". This file moves that one +// engine-facing read behind the facade so the route stops naming +// `lib/transcript.js` directly, and so the place where the enrichment +// can fail is stated once, in the engine layer, instead of being implied +// by control flow in a route. +// +// Why this family has a SOFT gate and the tree family has a HARD one. +// This is the one place where copying B1's shape verbatim would have +// been wrong, so the difference is deliberate and load-bearing: +// +// - #8 session-tree is 100% engine data. No session listing, no tree. +// Answering `501` is the only honest response, and the endpoint +// already had a documented "cannot read" shape to fall back on. +// - #11 export is mostly NOT engine data. A provider that declared +// `sessionCrud: none` would still leave the full user-visible chat +// exportable from `sessions.json`. Gating the endpoint hard would +// REMOVE working functionality in response to a declaration about a +// capability the endpoint does not actually depend on — and it would +// break the explicit "never block export" contract, which is this +// repository's #110 discipline applied in the other direction: a +// missing enrichment must not be dressed up as a failure, and a +// missing capability must not be dressed up as one either. +// +// So the gate here REPORTS and never throws. `checkSessionExportCapability` +// answers what the provider declared, and the read degrades through the +// endpoint's own pre-existing fail-soft channel +// (`_meta.mcode_unavailable` + `mcode_unavailable_reason`) rather than +// through an HTTP status. The 501 machinery in `errors.js` stays +// untouched and unused by this family — that is a policy statement, not +// an oversight, and the test suite pins it. +// +// What this file deliberately does NOT do: +// +// - It does not own the export. Parsing webui chat lines, merging the +// two message sources, and rendering md/json are the route's job and +// stay there; they are presentation, not engine access. Only the +// transcript read crosses this seam. +// - It does not own the session lookup. `_findSession` resolves a +// webui id or an `mvs_` id against `sessions.json` — webui's own +// store, governed by no engine capability. +// - It does not construct a host. +// +// Boot-path weight. `app.js` imports the routes, the routes import this +// file, so this file is on the boot path. It statically imports nothing +// heavier than `capabilities.js` and `index.js`; `lib/transcript.js` +// and `lib/config.js` are reached through `await import()` inside the +// functions — the M1 lesson again. +// +// Provider selection is M4's job, same as B1 and as the tree family. + +import { DEFAULT_ENGINE_PROVIDER_ID, getEngineProvider } from "./index.js"; + +/** + * Transport → registered engine provider id. Same shape and same + * rationale as `session-reads.js#providerByTransport` and + * `session-tree-reads.js#providerByTransport`; kept per-family so each + * family owns its own gate policy. Collapse the three in M4, not here. + * + * Built per call rather than frozen at module scope: `engine/index.js` + * re-exports this module, so a module-level table would read + * `DEFAULT_ENGINE_PROVIDER_ID` while that binding is still in its + * temporal dead zone on a cold `import("./engine/index.js")`. + * + * @returns {Readonly>} + */ +function providerByTransport() { + return Object.freeze({ runtime: DEFAULT_ENGINE_PROVIDER_ID }); +} + +/** + * The declaration this endpoint's ENRICHMENT needs. + * + * `sessionCrud` / `getSession` is the honest mapping — the same pair + * B1's `GET /api/acp-session-title` uses, because reading a session's + * transcript is reading that session. It is the ENRICHMENT that is + * declared, not the export: see the soft-gate rationale in the header. + * + * @type {Readonly>} + */ +export const SESSION_EXPORT_ENDPOINTS = Object.freeze({ + "GET /api/sessions/:id/export": { + capability: "sessionCrud", + subItem: "getSession", + enforcement: "soft", + }, +}); + +/** + * Resolve the provider that answers export enrichment on `transport`, + * or `null` when none is registered yet. + * + * @param {string} transport One of the `MCODE_WEBUI_TRANSPORT` values. + * @returns {{id: string, transport: string, capabilities: object}|null} + */ +export function resolveSessionExportProvider(transport) { + const providerId = providerByTransport()[transport]; + if (!providerId) return null; + return getEngineProvider(providerId); +} + +/** + * Read the declaration for this endpoint WITHOUT enforcing it. + * + * Returns a descriptor whose `gate` field says what happened: + * + * - `"checked"` — provider resolved, capability is `full`. + * - `"unregistered-transport"` — no provider claims this transport yet. + * - `"capability-absent"` — the provider WAS found and DOES declare the + * capability as `none` (or `partial` missing this sub-item). This is + * the branch that makes the soft gate visible: the caller's next + * move is to degrade the ENRICHMENT, not to fail the request. + * - `"partial"` — provider is `partial` and this sub-item is + * absent; the endpoint still degrades, but the descriptor says so + * precisely. + * + * Deliberately never throws `EngineCapabilityNotSupportedError`. A + * caller that wants the hard behaviour (the tree family) must ask for + * it explicitly; that asymmetry is the point of splitting the two. + * A genuinely unknown endpoint key is still a plain Error — caller + * confusion is not a capability question. + * + * @param {string} endpoint A key of SESSION_EXPORT_ENDPOINTS. + * @param {string} transport The active transport. + * @returns {{endpoint: string, gate: string, provider: string|null, capability: string|null, subItem: string|null, enforcement: "soft"}} + */ +export function checkSessionExportCapability(endpoint, transport) { + const need = SESSION_EXPORT_ENDPOINTS[endpoint]; + if (need === undefined) { + const err = new Error( + `checkSessionExportCapability: "${endpoint}" is not part of the session-export family ` + + `(known: ${Object.keys(SESSION_EXPORT_ENDPOINTS).join(", ")})`, + ); + err.code = "unknown_session_export_endpoint"; + throw err; + } + const base = { + endpoint, + provider: null, + capability: need.capability, + subItem: need.subItem, + enforcement: need.enforcement, + }; + const provider = resolveSessionExportProvider(transport); + if (!provider) return { ...base, gate: "unregistered-transport" }; + const entry = provider.capabilities ? provider.capabilities[need.capability] : undefined; + const descriptor = { ...base, provider: provider.id }; + if (entry && entry.level === "full") { + return { ...descriptor, gate: "checked" }; + } + if (entry && entry.level === "partial") { + const absent = Array.isArray(entry.missing) && entry.missing.includes(need.subItem); + return { ...descriptor, gate: absent ? "partial" : "checked" }; + } + // `none`, or no entry at all — the provider was found and does not + // offer this. Report it; the caller degrades the enrichment. + return { ...descriptor, gate: "capability-absent" }; +} + +/** + * Lazily resolve the transcript module and the active transport. + * Dynamic: `lib/transcript.js` reaches the sqlite resolver and the + * settings chain, neither of which may sit on the boot path. + */ +async function exportDeps() { + const [transcript, config] = await Promise.all([ + import("../lib/transcript.js"), + import("../lib/config.js"), + ]); + return { transcript, transport: config.MCODE_WEBUI_TRANSPORT }; +} + +/** + * Where the enrichment's bytes came from. `engine` when the transcript + * reader answered; `none` when it did not (and the caller degrades). + * + * @typedef {"engine" | "none"} SessionExportSource + */ + +/** + * The #11 engine-facing read: one session's transcript, best-effort. + * + * The returned `ok` / `reason` / `messages` are `readMcodeTranscript`'s + * own values, forwarded verbatim — this facade never invents a reason + * code and never converts a failure into an exception, because the + * endpoint's `_meta.mcode_unavailable` / `mcode_unavailable_reason` + * contract is built on those exact strings. `probeTable` and `probe` + * are the reader's own `source` / `probe`, renamed so they cannot be + * confused with this layer's `source`. + * + * Async even though the reader is synchronous (better-sqlite3 is sync): + * the route is already async, and a uniform awaitable `readEngine*` + * seam means a provider-backed transcript source that IS async (a + * network engine) needs no signature change at this layer. + * + * @param {object} [options] + * @param {string} [options.mcodeSessionId] The `mvs_…` id to read. + * @param {string} [options.endpoint] Endpoint key for the + * declaration check; defaults to `/api/sessions/:id/export`. + * @param {string} [options.transport] Transport override; defaults + * to the active `MCODE_WEBUI_TRANSPORT`. + * @returns {Promise<{mcodeSessionId: string, messages: Array, ok: boolean, reason: string|null, probeTable: string|null, probe: string|null, source: SessionExportSource, gate: object, transport: string}>} + */ +export async function readEngineSessionTranscript(options = {}) { + const endpoint = options.endpoint || "GET /api/sessions/:id/export"; + const deps = await exportDeps(); + const transport = options.transport || deps.transport; + const gate = checkSessionExportCapability(endpoint, transport); + const mcodeSessionId = options.mcodeSessionId || ""; + const r = deps.transcript.readMcodeTranscript(mcodeSessionId); + return { + mcodeSessionId, + messages: Array.isArray(r.messages) ? r.messages : [], + ok: r.ok === true, + reason: r.ok === true ? null : r.reason || "unknown", + probeTable: r.source || null, + probe: r.probe || null, + source: r.ok === true ? "engine" : "none", + gate, + transport, + }; +} diff --git a/packages/webui/server/engine/session-reads.js b/packages/webui/server/engine/session-reads.js new file mode 100644 index 00000000..b1df87d7 --- /dev/null +++ b/packages/webui/server/engine/session-reads.js @@ -0,0 +1,302 @@ +// webui/server/engine/session-reads.js +// +// Migration step M3, batch B1: the directory-read family (目录读族) — +// the five endpoints that only ever ASK the engine what it knows: +// +// #9 GET /api/acp-sessions — sidebar session list (cwd filtered) +// #10 GET /api/acp-session-title — one session's title +// #72 GET /api/protocol/list-sessions — remote-control session list (all) +// #74 GET /api/state — snapshot, mcodeSessions mirror source +// #75 GET /api/health — engine version, /api/state sibling +// +// What this file is for. Before M3 a route asked the ACP client directly +// and inherited whatever the transport happened to be. After M3 the route +// asks the facade, the facade checks the provider's DECLARATION first, and +// a provider that does not offer the read answers 501 through +// `invokeHandler`'s `EngineCapabilityNotSupportedError` mapping instead of +// quietly returning `[]` (the #110 fake-success failure mode). +// +// What this file deliberately does NOT do: +// +// - It does not re-implement listing. `lib/acp-client.js` already owns +// the transport switch and already normalises the catalogue host's +// `TuiSession` through `lib/catalogue-sessions.js#projectTuiSessionToAcp` +// — the one normalizer whose output is the ACP `session/list` wire +// shape the sidebar tree speaks. A second normalizer here would be a +// second answer to a shape question that must have exactly one. +// - It does not construct a host. `getCatalogueHost()` is the process +// singleton; this module only forwards to it (see `Never build a +// second host`). +// - It does not build a host-shaped error of its own for "the engine +// could not boot": a read family that fails over to the ACP mirror +// reports WHERE the bytes came from (`source`) instead of pretending +// the engine answered. Only endpoints whose whole contract is the +// engine (plugins / turn-diff) answer `RUNTIME_UNAVAILABLE`. +// +// Boot-path weight. `app.js` imports the routes, the routes import this +// file, so this file is on the boot path. It therefore statically imports +// nothing heavier than `capabilities.js` and `index.js` (both pure +// declaration modules); `lib/acp-client.js` and `lib/config.js` are +// reached through `await import()` inside the functions. That split is the +// M1 lesson — putting the `@mavis/*` tree on the boot path once cost +// 209ms → 2700ms of server start and broke the integration tests' 3s +// window. +// +// Provider selection is M4's job. `providerByTransport()` maps a transport +// to a REGISTERED provider id; today only `runtime` has one, so under the +// default `acp` transport there is no declaration to check and the gate +// reports `gate: "unregistered-transport"` instead of inventing one. When +// M4 registers the ACP provider this table gains its entry and the gate +// starts answering for the default transport too. + +import { assertEngineCapability } from "./capabilities.js"; +import { DEFAULT_ENGINE_PROVIDER_ID, getEngineProvider } from "./index.js"; + +/** + * Transport → registered engine provider id. Absent means "no provider + * claims this transport yet" (M4), NOT "the capability is unavailable" — + * the two answer differently on purpose: a missing provider answer is the + * pre-M4 passthrough, an unsupported capability answer is 501. + * + * Built per call rather than frozen at module scope: `engine/index.js` + * re-exports this module, so a module-level table would read + * `DEFAULT_ENGINE_PROVIDER_ID` while that binding is still in its temporal + * dead zone on a cold `import("./engine/index.js")` — the evaluation order + * of a re-export is the importer's, not this module's. Every consumer of + * the table is a function anyway. + * + * @returns {Readonly>} + */ +function providerByTransport() { + return Object.freeze({ runtime: DEFAULT_ENGINE_PROVIDER_ID }); +} + +/** + * The declaration each endpoint in this family needs, and the sub-item it + * needs from that capability. `subItem` is the provider method the route + * ultimately depends on, so a `partial` declaration that omits exactly that + * method yields 501 naming the method rather than a generic refusal. + * + * `/api/health` is `null`: reading the engine's own version is not any of + * the 14 matrix keys (`updateCheck` is about checking for a NEW version, + * not reporting the installed one), and inventing a key here would put a + * lie in the capability registry. See `readEngineVersion` for what the + * endpoint does instead. + * + * @type {Readonly>} + */ +export const SESSION_READ_ENDPOINTS = Object.freeze({ + "GET /api/acp-sessions": { capability: "sessionCrud", subItem: "listSessions" }, + "GET /api/acp-session-title": { capability: "sessionCrud", subItem: "getSession" }, + "GET /api/protocol/list-sessions": { capability: "sessionCrud", subItem: "listSessions" }, + "GET /api/state": { capability: "sessionCrud", subItem: "listSessions" }, + "GET /api/health": null, +}); + +/** + * Resolve the provider that answers catalogue reads on `transport`, or + * `null` when none is registered yet. + * + * @param {string} transport One of the `MCODE_WEBUI_TRANSPORT` values. + * @returns {{id: string, transport: string, capabilities: object}|null} + */ +export function resolveSessionReadProvider(transport) { + const providerId = providerByTransport()[transport]; + if (!providerId) return null; + return getEngineProvider(providerId); +} + +/** + * Check one endpoint of this family against the active provider's + * declaration. Throws `EngineCapabilityNotSupportedError` — which + * `app.js#invokeHandler` turns into 501 — when the declaration says the + * capability (or the exact sub-item) is absent. + * + * @param {string} endpoint A key of SESSION_READ_ENDPOINTS. + * @param {string} transport The active transport. + * @returns {{endpoint: string, gate: string, provider: string|null, capability: string|null, subItem: string|null}} + */ +export function assertSessionReadCapability(endpoint, transport) { + const need = SESSION_READ_ENDPOINTS[endpoint]; + if (need === undefined) { + // Caller confusion, not an engine limitation — a plain Error so the + // HTTP layer never answers 501 for a typo in webui's own code. + const err = new Error( + `assertSessionReadCapability: "${endpoint}" is not part of the session-read family ` + + `(known: ${Object.keys(SESSION_READ_ENDPOINTS).join(", ")})`, + ); + err.code = "unknown_session_read_endpoint"; + throw err; + } + const provider = resolveSessionReadProvider(transport); + if (need === null) { + return { + endpoint, + gate: "no-capability-key", + provider: provider ? provider.id : null, + capability: null, + subItem: null, + }; + } + if (!provider) { + return { + endpoint, + gate: "unregistered-transport", + provider: null, + capability: need.capability, + subItem: need.subItem, + }; + } + assertEngineCapability(provider.capabilities, need.capability, provider.id, need.subItem); + return { + endpoint, + gate: "checked", + provider: provider.id, + capability: need.capability, + subItem: need.subItem, + }; +} + +// --------------------------------------------------------------------------- +// Reads +// --------------------------------------------------------------------------- + +/** + * Where a list read's bytes actually came from. `catalogue` means the + * in-process host answered (and went through `catalogue-sessions.js`); + * `acp` means the `mcode acp` subprocess answered; `acp-fallback` means + * the transport ASKED for the catalogue host and the host was null, so the + * ACP mirror answered instead — reported rather than hidden, because a + * sidebar that silently loses its runtime path is exactly the degradation + * this batch exists to make visible. + * + * @typedef {"catalogue" | "acp" | "acp-fallback"} SessionReadSource + */ + +/** + * Lazily resolve the acp-client module and the active transport. Dynamic + * on both counts: `lib/acp-client.js` pulls `acp.mjs` and the settings + * chain, `lib/config.js` reads env — neither may sit on the boot path. + */ +async function readDeps() { + const [acp, config] = await Promise.all([ + import("../lib/acp-client.js"), + import("../lib/config.js"), + ]); + return { acp, transport: config.MCODE_WEBUI_TRANSPORT }; +} + +/** + * Did the catalogue host answer on this transport? Returns `false` when + * the transport never wanted the catalogue, and also when it wanted it but + * the host failed to boot (the `acp-fallback` case). + */ +async function catalogueAnswered(acp, transport) { + if (transport !== "runtime") return false; + const host = await acp.getCatalogueHost(); + return host !== null && host !== undefined; +} + +/** + * Every session the engine knows, across all workspaces — the #72 + * (`/api/protocol/list-sessions`) read. The caller applies its own cwd + * filter, exactly as the endpoint did before, so the filtering rule and + * the response shape stay in one place. + * + * @param {object} [options] + * @param {string} [options.endpoint] Endpoint key for the declaration + * check; defaults to `/api/protocol/list-sessions`. + * @returns {Promise<{sessions: Array, source: SessionReadSource, gate: object, transport: string}>} + */ +export async function readEngineSessionList(options = {}) { + const endpoint = options.endpoint || "GET /api/protocol/list-sessions"; + const { acp, transport } = await readDeps(); + const gate = assertSessionReadCapability(endpoint, transport); + const answered = await catalogueAnswered(acp, transport); + return { + sessions: await acp.listAllMcodeSessions(), + source: transport !== "runtime" ? "acp" : answered ? "catalogue" : "acp-fallback", + gate, + transport, + }; +} + +/** + * The workspace-filtered session list — the #9 (`/api/acp-sessions`) and + * #74 (`/api/state` `mcodeSessions` mirror) read. Same 30s cache and same + * path normalisation as before, because the call goes to the same + * `getMcodeSessionsForWorkspace`. + * + * @param {object} options + * @param {string} [options.cwd] Workspace to filter by; empty means + * "no filter" and the caller decides. + * @param {string} [options.endpoint] Endpoint key for the declaration + * check; defaults to `/api/acp-sessions`. + * @returns {Promise<{sessions: Array, source: SessionReadSource, gate: object, transport: string}>} + */ +export async function readEngineSessionListForWorkspace(options = {}) { + const endpoint = options.endpoint || "GET /api/acp-sessions"; + const { acp, transport } = await readDeps(); + const gate = assertSessionReadCapability(endpoint, transport); + const answered = await catalogueAnswered(acp, transport); + const sessions = await acp.getMcodeSessionsForWorkspace(options.cwd || ""); + return { + sessions, + source: transport !== "runtime" ? "acp" : answered ? "catalogue" : "acp-fallback", + gate, + transport, + }; +} + +/** + * One session's title — the #10 (`/api/acp-session-title`) read. `null` + * for "no such session" and `null` for "engine has no title", which is the + * contract the endpoint has always had; the bridge does not merge them. + * + * @param {object} options + * @param {string} options.sessionId + * @returns {Promise<{sessionId: string, title: string|null, source: SessionReadSource, gate: object, transport: string}>} + */ +export async function readEngineSessionTitle(options = {}) { + const { acp, transport } = await readDeps(); + const gate = assertSessionReadCapability("GET /api/acp-session-title", transport); + const answered = await catalogueAnswered(acp, transport); + const sessionId = options.sessionId || ""; + return { + sessionId, + title: sessionId ? await acp.getMcodeSessionTitle(sessionId) : null, + source: transport !== "runtime" ? "acp" : answered ? "catalogue" : "acp-fallback", + gate, + transport, + }; +} + +/** + * The engine's installed version — the #75 (`/api/health`) read. + * + * The one honest answer available today comes from the ACP `initialize` + * reply's `agentInfo.version`; the in-process catalogue host exposes no + * version accessor (its surface is `adapter` / `cliService` / `apiHost` / + * `controller` / `application` / `applications`, see + * `providers/local-runtime-v2.js`), so "read it from the v2 host" as the + * batch plan imagined is not implementable without inventing a method. + * Rather than fabricate one, the bridge names the source it used and + * keeps the endpoint's `"unknown"` fallback for "nothing has attached + * yet". The shape of `/api/health` is untouched. + * + * @returns {Promise<{version: string, source: SessionReadSource, transport: string, gate: object}>} + */ +export async function readEngineVersion() { + const { acp, transport } = await readDeps(); + const gate = assertSessionReadCapability("GET /api/health", transport); + const info = acp.getMcodeServerInfo(); + return { + version: (info && info.version) || "unknown", + // The version is a protocol fact, not a catalogue fact: it is answered + // from the ACP `initialize` mirror under every transport, including + // `runtime`, where the mirror is simply empty until something attaches. + source: "acp", + transport, + gate, + }; +} diff --git a/packages/webui/server/engine/session-tree-reads.js b/packages/webui/server/engine/session-tree-reads.js new file mode 100644 index 00000000..cfb3a4e0 --- /dev/null +++ b/packages/webui/server/engine/session-tree-reads.js @@ -0,0 +1,220 @@ +// webui/server/engine/session-tree-reads.js +// +// Migration step M3, batch B2: the session-tree read (树读) — +// +// #8 GET /api/session-tree — the sidebar's Project → directory → +// session → subagent tree +// +// What this file is for. #8 is the main↔subagent communication spine: the +// hierarchy the user navigates is built from `parent_session_id`, and a +// child that fails to attach to its parent is a subagent the user cannot +// see. So the endpoint's payload is a frontend contract of the strictest +// kind here, and the route now reaches the engine through this file +// instead of calling `lib/session-tree.js` on its own: the facade checks +// the provider's DECLARATION first, then forwards to the one existing +// implementation. A provider that does not offer session listing answers +// 501 through `app.js#invokeHandler`'s EngineCapabilityNotSupportedError +// mapping rather than an empty tree, which the sidebar would render as +// "this project has no sessions" (#110 fake-success failure mode). +// +// What this file deliberately does NOT do: +// +// - It does not re-assemble the tree. `lib/session-tree.js` owns the +// level mapping (project / directory / branch / subagent), the +// git-based project resolution and the 15s cache. A second +// assembler here would be a second answer to a hierarchy question +// that must have exactly one — and getting it wrong by one level is +// precisely the failure this batch is gated on. +// - It does not normalise nodes. The node shape +// `{id, title, agent, kind, status, updatedAt, children}` is +// whatever `buildTree` produces, byte-for-byte. Note what is NOT in +// it: the response carries NO `parent_session_id` key. The hierarchy +// is expressed structurally through `children`; `parent_session_id` +// exists only inside the db read. `test/lib/engine/session-tree-reads.test.js` +// pins the exact key set so a future "helpful" addition is caught. +// - It does not construct a host. The tree is assembled from the +// engine's own runtime db (see the source note below), so there is +// no host in this path at all — see `Never build a second host`. +// - It does not widen the engine's own degradation. A db that cannot +// be read is still `{ok:false, reason}` with HTTP 200, exactly as +// before; that is the sidebar's documented fallback to the wrapper +// list, and the capability gate is a different question (may this +// provider list sessions AT ALL) from "could we read the db right +// now" (could we read it THIS TIME). +// +// Transport. Unlike the B1 read family, this read is deliberately +// transport-independent: it reads `local_runtime_sessions` in the +// runtime db, the engine's own persistent store, which both the +// `runtime` and the `acp` transport can see. `source` therefore +// reports `runtime-db` under every transport rather than pretending +// to be a catalogue answer. The DECLARATION check is still +// transport-keyed, because which provider is active is a transport +// question even when the read itself is not. +// +// Boot-path weight. `app.js` imports the routes, the routes import this +// file, so this file is on the boot path. It therefore statically +// imports nothing heavier than `capabilities.js` and `index.js` (both +// pure declaration modules); `lib/session-tree.js` and `lib/config.js` +// are reached through `await import()` inside the functions. That split +// is the M1 lesson — putting the `@mavis/*` tree on the boot path once +// cost 209ms → 2700ms of server start and broke the integration tests' +// 3s window. +// +// Provider selection is M4's job, same as B1: `providerByTransport()` +// maps a transport to a REGISTERED provider id; today only `runtime` has +// one, so under the default `acp` transport the gate reports +// `gate: "unregistered-transport"` instead of inventing one. + +import { assertEngineCapability } from "./capabilities.js"; +import { DEFAULT_ENGINE_PROVIDER_ID, getEngineProvider } from "./index.js"; + +/** + * Transport → registered engine provider id. Absent means "no provider + * claims this transport yet" (M4), NOT "the capability is unavailable" — + * the two answer differently on purpose, exactly as in + * `session-reads.js#providerByTransport`, which this mirrors rather than + * merges: the two families have separate gate semantics (see + * `session-export.js` for the soft-gate counterpart) and a shared table + * would force one of them to inherit the other's policy. + * + * Built per call rather than frozen at module scope: `engine/index.js` + * re-exports this module, so a module-level table would read + * `DEFAULT_ENGINE_PROVIDER_ID` while that binding is still in its + * temporal dead zone on a cold `import("./engine/index.js")`. Every + * consumer of the table is a function anyway. + * + * @returns {Readonly>} + */ +function providerByTransport() { + return Object.freeze({ runtime: DEFAULT_ENGINE_PROVIDER_ID }); +} + +/** + * The declaration this endpoint needs, and the sub-item it needs from + * that capability. + * + * `sessionCrud` / `listSessions` is the honest mapping, and it is the + * same pair B1's `GET /api/protocol/list-sessions` uses: both endpoints + * answer "every session the engine knows, across all workspaces", and + * the tree is that list plus a hierarchy. The tree additionally needs + * the `parent_session_id` column, but that is not a separate + * provider method — it is a column of the same rows, so naming a + * sub-item that no provider enumerates would be a lie in the registry. + * + * @type {Readonly>} + */ +export const SESSION_TREE_ENDPOINTS = Object.freeze({ + "GET /api/session-tree": { capability: "sessionCrud", subItem: "listSessions" }, +}); + +/** + * Resolve the provider that answers the tree read on `transport`, or + * `null` when none is registered yet. + * + * @param {string} transport One of the `MCODE_WEBUI_TRANSPORT` values. + * @returns {{id: string, transport: string, capabilities: object}|null} + */ +export function resolveSessionTreeProvider(transport) { + const providerId = providerByTransport()[transport]; + if (!providerId) return null; + return getEngineProvider(providerId); +} + +/** + * Check the tree read against the active provider's declaration. Throws + * `EngineCapabilityNotSupportedError` — which `app.js#invokeHandler` + * turns into 501 — when the declaration says the capability (or the + * exact sub-item) is absent. + * + * @param {string} endpoint A key of SESSION_TREE_ENDPOINTS. + * @param {string} transport The active transport. + * @returns {{endpoint: string, gate: string, provider: string|null, capability: string|null, subItem: string|null}} + */ +export function assertSessionTreeCapability(endpoint, transport) { + const need = SESSION_TREE_ENDPOINTS[endpoint]; + if (need === undefined) { + // Caller confusion, not an engine limitation — a plain Error so the + // HTTP layer never answers 501 for a typo in webui's own code. + const err = new Error( + `assertSessionTreeCapability: "${endpoint}" is not part of the session-tree family ` + + `(known: ${Object.keys(SESSION_TREE_ENDPOINTS).join(", ")})`, + ); + err.code = "unknown_session_tree_endpoint"; + throw err; + } + const provider = resolveSessionTreeProvider(transport); + if (!provider) { + return { + endpoint, + gate: "unregistered-transport", + provider: null, + capability: need.capability, + subItem: need.subItem, + }; + } + assertEngineCapability(provider.capabilities, need.capability, provider.id, need.subItem); + return { + endpoint, + gate: "checked", + provider: provider.id, + capability: need.capability, + subItem: need.subItem, + }; +} + +/** + * Lazily resolve the session-tree module and the active transport. + * Dynamic on both counts: `lib/session-tree.js` reaches the sqlite + * resolver and the settings chain, `lib/config.js` reads env — neither + * may sit on the boot path. + */ +async function treeDeps() { + const [tree, config] = await Promise.all([ + import("../lib/session-tree.js"), + import("../lib/config.js"), + ]); + return { tree, transport: config.MCODE_WEBUI_TRANSPORT }; +} + +/** + * Where the tree's bytes came from. Always `runtime-db`: the tree is + * assembled from `local_runtime_sessions` in the engine's own runtime + * db, which is not a transport-switched surface (see the transport note + * in the file header). The value exists so a consumer never has to + * guess whether the ACP mirror answered instead. + * + * @typedef {"runtime-db"} SessionTreeSource + */ + +/** + * The #8 (`GET /api/session-tree`) read. + * + * Forwards `options` straight to `getSessionTree`, so `force` keeps its + * meaning (`?refresh=1` bypasses the 15s cache) and the `cached` field + * keeps its shape. The returned `tree` is the endpoint's payload + * verbatim — including the `ok:false` / `reason` soft-fail shape for a + * missing or unreadable db, which this facade deliberately does not + * convert into an error. + * + * @param {object} [options] + * @param {boolean} [options.force] Bypass the 15s cache. + * @param {string} [options.endpoint] Endpoint key for the declaration + * check; defaults to `/api/session-tree`. + * @param {string} [options.now] Clock injection, forwarded as-is. + * @param {string} [options.transport] Transport override; defaults to the + * active `MCODE_WEBUI_TRANSPORT`. Exists so tests can exercise + * both the `runtime` and the unregistered `acp` branch without + * mutating process env. + * @returns {Promise<{tree: object, source: SessionTreeSource, gate: object, transport: string}>} + */ +export async function readEngineSessionTree(options = {}) { + const endpoint = options.endpoint || "GET /api/session-tree"; + const deps = await treeDeps(); + const transport = options.transport || deps.transport; + const gate = assertSessionTreeCapability(endpoint, transport); + const tree = deps.tree.getSessionTree({ + force: options.force === true, + ...(options.now === undefined ? {} : { now: options.now }), + }); + return { tree, source: "runtime-db", gate, transport }; +} diff --git a/packages/webui/server/engine/usage-reads.js b/packages/webui/server/engine/usage-reads.js new file mode 100644 index 00000000..fa0b774b --- /dev/null +++ b/packages/webui/server/engine/usage-reads.js @@ -0,0 +1,434 @@ +// webui/server/engine/usage-reads.js +// +// Migration step M3, batch B3: the usage family (用量族) — the four +// endpoints that answer "how much has this cost, and when will it run +// out": +// +// #15 POST /api/usage — plan quota (5h / weekly) read +// #16 POST /api/usage-trigger — the same read, recorded as a sample +// #17 GET /api/usage-real — real per-session token usage +// #19 GET /api/usage/forecast — quota-exhaustion prediction +// +// What this file is for. Three of these four numbers decide what a user +// does next — refresh, switch model, stop working — and each of them is a +// DERIVED quantity, not a counter. #15/#16 re-derive the plan windows out +// of the engine's account projection. #17 re-derives `contextUsed` out of +// three separate token totals. #19 re-derives an exhaustion time out of a +// least-squares fit. A refactor that "cleans up" one of those formulas +// changes what the user sees and reports nothing, which is the failure +// mode this batch is gated on. So the derivations live HERE, once, named, +// and tested on their inputs — the route only assembles JSON. +// +// What this file deliberately does NOT do: +// +// - It does not re-read the database. `lib/mavis-usage.js` owns the SQL, +// the `node:sqlite` / `sqlite3`-spawn dual path and the NULL-to-zero +// coercion; `lib/usage.js` owns the quota-window copy into `cs.usage`; +// `lib/quota-forecast.js` owns the NDJSON history and the least-squares +// fit. A second reader over `local_runtime_token_usage` would be a +// second answer to "what did this session cost". +// - It does not construct a host. #17's data currently comes from the +// engine's own SQLite file, not from a live `CliService` — see the +// `getSessionUsage` note below for why the provider call is deferred, +// and for what would have to be true before it is not. +// - It does not widen the engine's own degradation. `getMavisTokenUsage` +// returns `null` for "no such session / no db / no rows", and the route +// turns that into `{ok:true, found:false, …}` with HTTP 200. That +// answer is the endpoint's long-standing contract and it is a +// different question from "may this provider report usage at all". +// +// The `getSessionUsage` question, stated once because it is the batch's +// most load-bearing decision. The v2 provider declares `usageStats: full`, +// and `CliService#getSessionUsage` is a real method that reads the SAME +// `local_runtime_token_usage` table this endpoint already reads. Routing +// through it anyway today would be a behaviour change dressed as a +// refactor, for three measured reasons: +// +// 1. It only exists under the `runtime` transport. The catalogue host is +// booted by `acp-client.js#transportWantsCatalogue()`, which is +// `MCODE_WEBUI_TRANSPORT === "runtime"`. The DEFAULT transport is +// `acp` (`lib/config.js`), and under it there is no `CliService` to +// call — so the switch would take the endpoint from "always answers" +// to "answers on one opt-in transport". +// 2. Its shape is not this endpoint's shape. `getSessionUsage` answers +// `{summary, rows: UsageView[]}`; the endpoint answers a per-column +// aggregate plus `rows` as a COUNT. Rebuilding the aggregate from +// `rows` would re-derive `totalReasoning` and `contextUsed` from a +// different starting point — exactly the silent numeric drift this +// batch forbids. +// 3. It would put the v2 TypeScript dependency tree on the answer path +// of an endpoint that currently needs nothing from it (the M1 lesson). +// +// So the declaration names the provider method the endpoint DEPENDS ON — +// which is what a declaration is for — and the read keeps using the file +// the provider itself would read. M4 is where the two are allowed to meet. +// +// Boot-path weight. `app.js` imports the routes, the routes import this +// file, so this file is on the boot path. It therefore statically imports +// nothing heavier than `capabilities.js` and `index.js` (both pure +// declaration modules); `lib/usage.js`, `lib/mavis-usage.js`, +// `lib/quota-forecast.js` and `lib/config.js` are reached through +// `await import()` inside the functions. That split is the M1 lesson — +// putting the `@mavis/*` tree on the boot path once cost 209ms → 2700ms of +// server start and broke the integration tests' 3s window. +// +// Provider selection is M4's job, same as B1 and B2: `providerByTransport()` +// maps a transport to a REGISTERED provider id; today only `runtime` has +// one, so under the default `acp` transport the gate reports +// `gate: "unregistered-transport"` instead of inventing one. + +// `node:fs` is a builtin, not a project dependency: the boot-path promise +// below is about not dragging lib/ or @mavis/* trees in, and this costs +// nothing. It is here for one caller — #17's `dbExists`, which the route +// used to compute itself from a path constant it imported at module scope. +import { existsSync } from "node:fs"; + +import { assertEngineCapability } from "./capabilities.js"; +import { DEFAULT_ENGINE_PROVIDER_ID, getEngineProvider } from "./index.js"; + +/** + * Transport → registered engine provider id. Absent means "no provider + * claims this transport yet" (M4), NOT "the capability is unavailable" — + * the two answer differently on purpose, exactly as in + * `session-reads.js#providerByTransport` and + * `session-tree-reads.js#providerByTransport`, which this mirrors rather + * than merges: the four families have separate read contracts and a shared + * table would force one of them to inherit another's policy. + * + * Built per call rather than frozen at module scope: `engine/index.js` + * re-exports this module, so a module-level table would read + * `DEFAULT_ENGINE_PROVIDER_ID` while that binding is still in its temporal + * dead zone on a cold `import("./engine/index.js")`. Every consumer of the + * table is a function anyway. + * + * @returns {Readonly>} + */ +function providerByTransport() { + return Object.freeze({ runtime: DEFAULT_ENGINE_PROVIDER_ID }); +} + +/** + * The declaration each endpoint of this family needs, and the sub-item it + * needs from that capability. + * + * - #15 / #16 are `authCredentials` / `getAccountStatus`. The plan tier + * and both window percentages come from the engine's account + * projection, and the declaration names that method explicitly ("the + * engine holds the credential, so it is the only side that may call + * MiniMax's quota endpoint" — see `lib/usage.js`). A `partial` that + * dropped exactly `getAccountStatus` would answer 501 naming it rather + * than a generic refusal. + * - #17 is `usageStats` / `getSessionUsage` — the same pair the v2 + * declaration enumerates under `usageStats`. See the header for why + * the read does not yet call that method. + * - #19 is `null`, and this is the one row a reader will double-take. + * The forecast reads `~/.mcode-webui/usage-history.ndjson`, a file + * webui itself appends to; it calls no engine surface at all. The + * numbers in it ORIGINATED in the engine, but a read that touches no + * engine surface must not be gated on an engine capability — that is + * the same lie B1 declined for `/api/health`, and gating it hard would + * remove a working endpoint in response to a declaration about + * something it does not depend on. The precedent for a soft family + * that DOES cross the seam is B2's export enrichment + * (`_meta.mcode_unavailable`); #19 needs none of that, because there is + * no enrichment to lose. + * + * @type {Readonly>} + */ +export const USAGE_READ_ENDPOINTS = Object.freeze({ + "POST /api/usage": { capability: "authCredentials", subItem: "getAccountStatus" }, + "POST /api/usage-trigger": { capability: "authCredentials", subItem: "getAccountStatus" }, + "GET /api/usage-real": { capability: "usageStats", subItem: "getSessionUsage" }, + "GET /api/usage/forecast": null, +}); + +/** + * Resolve the provider that answers usage reads on `transport`, or `null` + * when none is registered yet. + * + * @param {string} transport One of the `MCODE_WEBUI_TRANSPORT` values. + * @returns {{id: string, transport: string, capabilities: object}|null} + */ +export function resolveUsageReadProvider(transport) { + const providerId = providerByTransport()[transport]; + if (!providerId) return null; + return getEngineProvider(providerId); +} + +/** + * Check one endpoint of this family against the active provider's + * declaration. Throws `EngineCapabilityNotSupportedError` — which + * `app.js#invokeHandler` turns into 501 — when the declaration says the + * capability (or the exact sub-item) is absent. + * + * @param {string} endpoint A key of USAGE_READ_ENDPOINTS. + * @param {string} transport The active transport. + * @returns {{endpoint: string, gate: string, provider: string|null, capability: string|null, subItem: string|null}} + */ +export function assertUsageReadCapability(endpoint, transport) { + const need = USAGE_READ_ENDPOINTS[endpoint]; + if (need === undefined) { + // Caller confusion, not an engine limitation — a plain Error so the + // HTTP layer never answers 501 for a typo in webui's own code. + const err = new Error( + `assertUsageReadCapability: "${endpoint}" is not part of the usage family ` + + `(known: ${Object.keys(USAGE_READ_ENDPOINTS).join(", ")})`, + ); + err.code = "unknown_usage_read_endpoint"; + throw err; + } + const provider = resolveUsageReadProvider(transport); + if (need === null) { + return { + endpoint, + gate: "no-capability-key", + provider: provider ? provider.id : null, + capability: null, + subItem: null, + }; + } + if (!provider) { + return { + endpoint, + gate: "unregistered-transport", + provider: null, + capability: need.capability, + subItem: need.subItem, + }; + } + assertEngineCapability(provider.capabilities, need.capability, provider.id, need.subItem); + return { + endpoint, + gate: "checked", + provider: provider.id, + capability: need.capability, + subItem: need.subItem, + }; +} + +// --------------------------------------------------------------------------- +// The derivations. Pure functions, exported, and tested on their INPUTS. +// --------------------------------------------------------------------------- + +/** + * The `contextUsed` figure `GET /api/usage-real` reports. + * + * CUMULATIVE input + output + reasoning, and deliberately NOT the + * per-turn figure. The two coexist in this repository and confusing them + * is the single most likely way for this endpoint to start lying: + * + * - `lib/mavis-usage.js#_buildUsageResult` publishes + * `lastTurnContextTokens` (last input + output + reasoning) and the + * chat flow stores it as `cs.context.tokens` — the CONTEXT BAR. One + * turn's worth, always ≤ the model's context limit. + * - `GET /api/usage-real` reports the session's CUMULATIVE spend + * (`v0.5.bx-10` fix: "context 实际是 input + output + reasoning"), + * which is why a 13-turn session can show 566k there. That is what + * the number has always meant on this endpoint and the frontend reads + * it as such. + * + * `cacheRead` / `cacheWrite` are excluded: they are a SUBSET of `input` + * (counted again by the engine inside the prompt), so adding them + * double-counts. `totalCacheWrite` is excluded for the same reason plus + * the fact that it is not part of the context window at all. + * + * Written as one expression, in the order the endpoint has always summed, + * over the SAME three fields the endpoint has always summed. That is + * deliberate: every field here arrives already coerced to a number by + * `_buildUsageResult` (`Number(x) || 0`), so no rounding point is + * introduced, and a `null` from a future provider coerces exactly the way + * the pre-facade expression coerced it. `test/lib/engine/usage-reads.test.js` + * pins the inputs, not just this number. + * + * @param {object} usage A `getMavisTokenUsage` result. + * @returns {number} + */ +export function contextUsedTokens(usage) { + return usage.totalInput + usage.totalOutput + usage.totalReasoning; +} + +// --------------------------------------------------------------------------- +// Reads +// --------------------------------------------------------------------------- + +/** + * Where each read's bytes actually came from. Three distinct producers, + * named rather than assumed: + * + * - `"account-status"` — the engine's `mcode/account/status` extension + * method, via `lib/mcode-rpc.js#getAccountStatus`. The same answer + * `quotaSnapshot` has always labelled `source: "acp"` in its own + * payload; the facade names the producer rather than the wire. + * - `"runtime-db"` — the engine's own `local_runtime_token_usage` table + * in its runtime sqlite, via `lib/mavis-usage.js`. Same vocabulary as + * B2's session tree: not a transport-switched surface. + * - `"history-file"` — webui's OWN `usage-history.ndjson`. The forecast + * read touches no engine surface, which is why its declaration row is + * `null`; this value keeps that honest at the call site. + * + * @typedef {"account-status" | "runtime-db" | "history-file"} UsageReadSource + */ + +/** + * #15 / #16 — the plan-quota read. + * + * `runUsageQuery` is the whole contract and is forwarded verbatim: it + * copies the engine's projection into `cs.usage`, appends at most one + * NDJSON history sample when `record` is on, pushes state, and returns the + * popover payload. The route writes that payload as the response body + * byte-for-byte, including its `ok:false` / `error` shape for an engine + * that could not be reached — the request itself succeeded, so the status + * stays 200. + * + * `record` is the difference between reading and measuring, and it is NOT + * defaulted here: `lib/usage.js` owns that default (`true`, the + * historical "a read is also a measurement" behaviour). The route passes + * the client's explicit `record !== false` through unchanged. + * + * @param {object} options + * @param {object} options.cs The webui client state `cs.usage` is written into. + * @param {string} options.cid Client id, for the state push. + * @param {boolean} [options.record] Append a forecast sample; see above. + * @param {string} [options.endpoint] Endpoint key for the declaration + * check; defaults to `/api/usage`. + * @param {string} [options.transport] Transport override; defaults to the + * active `MCODE_WEBUI_TRANSPORT`. Exists so tests can exercise both + * the `runtime` and the unregistered `acp` branch without mutating + * process env. + * @returns {Promise<{payload: object, source: UsageReadSource, gate: object, transport: string}>} + */ +export async function readEngineAccountQuota(options = {}) { + const endpoint = options.endpoint || "POST /api/usage"; + const usage = await import("../lib/usage.js"); + const config = await import("../lib/config.js"); + const transport = options.transport || config.MCODE_WEBUI_TRANSPORT; + const gate = assertUsageReadCapability(endpoint, transport); + const payload = await usage.runUsageQuery(options.cs, options.cid, { + record: options.record !== false, + }); + return { payload, source: "account-status", gate, transport }; +} + +/** + * #17 — the real per-session token usage. + * + * `usage` is `getMavisTokenUsage`'s own object, forwarded field for + * field: `rows`, the five totals, `firstTs`, `lastTs`, and the per-turn + * and cache-hit figures the chat flow also consumes. The facade adds + * exactly one derived number, `contextUsed` (see `contextUsedTokens`), and + * nothing else — in particular it does not re-derive `totalReasoning`, + * which is the database's own `SUM(reasoning_tokens)` and has exactly one + * correct source. + * + * `found:false` carries the same two facts the endpoint has always + * reported for "no session id yet / no such session": which database it + * looked in, and whether that database exists. `dbExists` is the + * `existsSync` the route used to do itself, moved behind the lazy + * `lib/config.js` boundary so `routes/usage.js` no longer names a path + * constant at module scope. + * + * @param {object} [options] + * @param {string|null} [options.mcodeSessionId] The `mvs_…` id to read. + * @param {string} [options.endpoint] Endpoint key for the declaration + * check; defaults to `/api/usage-real`. + * @param {string} [options.transport] Transport override; defaults to the + * active `MCODE_WEBUI_TRANSPORT`. + * @returns {Promise<{mcodeSessionId: string, found: boolean, usage: object|null, contextUsed: number|null, model: string|null, dbPath: string, dbExists: boolean, source: UsageReadSource, gate: object, transport: string}>} + */ +export async function readEngineSessionUsage(options = {}) { + const endpoint = options.endpoint || "GET /api/usage-real"; + const [mavis, config] = await Promise.all([ + import("../lib/mavis-usage.js"), + import("../lib/config.js"), + ]); + const transport = options.transport || config.MCODE_WEBUI_TRANSPORT; + const gate = assertUsageReadCapability(endpoint, transport); + const mcodeSessionId = options.mcodeSessionId || ""; + const dbPath = config.MAVIS_DB_PATH; + const dbExists = existsSync(dbPath); + const usage = mcodeSessionId ? await mavis.getMavisTokenUsage(mcodeSessionId) : null; + if (!usage) { + return { + mcodeSessionId, + found: false, + usage: null, + contextUsed: null, + model: null, + dbPath, + dbExists, + source: "runtime-db", + gate, + transport, + }; + } + // Best-effort and in that order: the endpoint has always answered even + // when the model lookup fails, and `getMavisTokenUsageModel` returns + // `null` for its own reasons (no db, no row, a `model` column that is + // NULL or empty). `(m && m.model) || null` is the endpoint's own + // fallback, kept verbatim. + const model = await mavis.getMavisTokenUsageModel(mcodeSessionId).catch(() => null); + return { + mcodeSessionId, + found: true, + usage, + contextUsed: contextUsedTokens(usage), + model: (model && model.model) || null, + dbPath, + dbExists, + source: "runtime-db", + gate, + transport, + }; +} + +/** + * #19 — the quota-exhaustion forecast. + * + * `readHistory` and `forecastExhaustion` are forwarded verbatim, which is + * what keeps the SEQUENCE continuous: the forecast for a given history + * prefix is a pure function of that prefix, and a refactor that re-read, + * re-filtered, re-sorted or re-sampled the history would shift every + * point of the curve without changing any single call's shape. + * `test/lib/engine/usage-reads.test.js#forecast sequence` pins the prefix + * series against the pre-refactor computation. + * + * The `try/catch` around `readHistory` is the endpoint's own belt-and- + * braces guard (the module already swallows FS errors; the catch is so a + * buggy extension can never break the endpoint) and it MOVES here with + * the read, because the read is what can fail. On failure the history is + * `[]`, and `forecastExhaustion([])` answers `reason: "no_history"` — + * byte-identical to the pre-facade body, which the UI renders as + * "collecting data…". + * + * @param {object} [options] + * @param {string} [options.endpoint] Endpoint key for the declaration + * check; defaults to `/api/usage/forecast`. + * @param {string} [options.transport] Transport override; defaults to the + * active `MCODE_WEBUI_TRANSPORT`. + * @param {object} [options.forecastOptions] Forwarded to + * `forecastExhaustion` (`minSamples`, `nowMs`); the endpoint passes + * neither today, and the defaults must stay the module's. + * @returns {Promise<{forecast: object, historyLength: number, source: UsageReadSource, gate: object, transport: string}>} + */ +export async function readEngineQuotaForecast(options = {}) { + const endpoint = options.endpoint || "GET /api/usage/forecast"; + const [quota, config] = await Promise.all([ + import("../lib/quota-forecast.js"), + import("../lib/config.js"), + ]); + const transport = options.transport || config.MCODE_WEBUI_TRANSPORT; + const gate = assertUsageReadCapability(endpoint, transport); + let history = []; + try { + history = quota.readHistory(); + } catch { + history = []; + } + return { + forecast: quota.forecastExhaustion(history, options.forecastOptions || {}), + historyLength: history.length, + source: "history-file", + gate, + transport, + }; +} diff --git a/packages/webui/server/lib/acp-client.js b/packages/webui/server/lib/acp-client.js index 68a36d3c..1d48b802 100644 --- a/packages/webui/server/lib/acp-client.js +++ b/packages/webui/server/lib/acp-client.js @@ -36,14 +36,22 @@ let _catalogueHostInitPromise = null; /** * The process-lifetime catalogue host singleton, booted on first call. * - * Exported because `/api/plugins/*` (routes/plugins.js) needs the runtime's - * `cliService` as its only data source, and the host is the single owner of - * that service. Routing plugins through the exported getter is deliberate: - * `transportWantsCatalogue()` below gates *session-list* traffic only — in - * ACP protocol there is no plugin method at all, so gating plugins on the - * transport would leave the panel dead in the default `acp` mode. Callers - * must never construct a second host: two CliService instances on one dataDir - * is both wasteful and a split-brain against the plugin/local-disable tables. + * Exported because `/api/plugins/*` (routes/plugins.js) and + * `/api/turn-diff*` (routes/turn-diff.js) need the runtime's `cliService` + * and `applications.session.diff` as their only data sources, and the host + * is the single owner of both. Since migration step M3's first batch (B0) + * those routes no longer import this module: they call the facade's + * `getEngineCatalogueHost()` (server/engine/host.js), which forwards here + * through a dynamic import, because `app.js` loads the engine facade at + * boot and this module carries the ACP client tree. The reasons below are + * the facade's reasons now, and the facade forwards them unchanged. + * + * Routing plugins through the host is deliberate: `transportWantsCatalogue()` + * below gates *session-list* traffic only — in ACP protocol there is no + * plugin method at all, so gating plugins on the transport would leave the + * panel dead in the default `acp` mode. Callers must never construct a + * second host: two CliService instances on one dataDir is both wasteful and + * a split-brain against the plugin/local-disable tables. * * Resolves to `null` when the runtime fails to boot; callers answer * `RUNTIME_UNAVAILABLE` rather than falling back to another path. diff --git a/packages/webui/server/routes/export.js b/packages/webui/server/routes/export.js index f1a5b5fb..fb5baf4c 100644 --- a/packages/webui/server/routes/export.js +++ b/packages/webui/server/routes/export.js @@ -25,14 +25,18 @@ // - ?download=true → Content-Disposition: attachment; filename="-." import { loadSessions } from "../lib/sessions.js"; -// v2 (2026-09-20 webui-manual-audit): _readMcodeTranscript's core moved to +// v2 (2026-09-20 webui-manual-audit): the transcript read moved to // lib/transcript.js so POST /api/sessions/switch can share the exact same // table-probing + fail-soft logic. Default probe set there is the legacy // 3-candidate list carried over VERBATIM (same SQL, same row mapping, same // reason strings) — export behavior is unchanged. existsSync / // MCODE_RUNTIME_DB / getMcodeBetterSqlite3 are no longer imported here // because only the extracted reader used them. -import { readMcodeTranscript } from "../lib/transcript.js"; +// +// M3-B2: this route no longer names that reader at all. The engine-facing +// half of the export goes through engine/session-export.js, which owns the +// soft gate and forwards the same `readMcodeTranscript` values verbatim. +import { readEngineSessionTranscript } from "../engine/session-export.js"; import { authorize } from "../lib/authorize.js"; import { pushAlert } from "../lib/alerts.js"; import { @@ -239,15 +243,6 @@ function _parseChatLines(lines) { }); } -// Best-effort: read mcode session transcript from runtime-state.sqlite. -// Returns { messages, ok } — ok=false means we set _meta.mcode_unavailable. -// v2 (2026-09-20 webui-manual-audit): body extracted to lib/transcript.js -// (readMcodeTranscript) — legacy probe set only, so this stays a pass-through -// and export behavior is byte-identical to the inline version. -function _readMcodeTranscript(mcodeSid) { - return readMcodeTranscript(mcodeSid); -} - // Merge webui messages + mcode transcript. Strategy: webui is authoritative // for the user-visible chat; mcode is best-effort enrichment (token usage, // full tool call payloads). When both exist for the same turn, mcode wins @@ -382,12 +377,19 @@ export async function handleExport(req, res, ctx) { // Parse webui chat → structured messages const webuiMsgs = _parseChatLines(Array.isArray(session.chat) ? session.chat : []); - // Best-effort mcode enrichment + // Best-effort mcode enrichment. + // + // M3-B2: the read goes through the engine facade, which reports the + // provider's declaration instead of enforcing it — export's primary + // source is `sessions.json`, not the engine, so a provider that cannot + // serve a transcript degrades THIS enrichment and nothing else. That is + // the "never block export" contract, kept verbatim: the `_meta` keys, + // the reason strings and the merged output are all unchanged. let mcodeMsgs = []; let mcodeUnavailable = false; let mcodeUnavailableReason = null; if (session.mcodeSessionId) { - const r = _readMcodeTranscript(session.mcodeSessionId); + const r = await readEngineSessionTranscript({ mcodeSessionId: session.mcodeSessionId }); if (r.ok) { mcodeMsgs = r.messages; } else { diff --git a/packages/webui/server/routes/health.js b/packages/webui/server/routes/health.js index a782b837..9da6f65e 100644 --- a/packages/webui/server/routes/health.js +++ b/packages/webui/server/routes/health.js @@ -8,22 +8,16 @@ import { DEFAULT_MODEL, DEFAULT_WORKSPACE, } from "../lib/config.js"; -import { getMcodeServerInfo } from "../lib/acp-client.js"; +// M3-B1 (engine facade): `mcodeVersion` is read through the facade so +// the endpoint records WHICH source answered. See +// `readEngineVersion` for why the answer is still the ACP `initialize` +// mirror — the in-process catalogue host exposes no version accessor, and +// inventing one is exactly the "claim a capability that does not exist" +// this batch exists to prevent. +import { readEngineVersion } from "../engine/session-reads.js"; -/** - * The engine's own version, from the `agentInfo` in its ACP `initialize` reply. - * - * This used to be a pinned constant, which meant the endpoint reported whatever - * version webui was written against rather than the one installed. Before a - * client attaches there is no version to report, hence `unknown` — the same - * value `/api/protocol/capabilities` uses for the same fact. - */ -function engineVersion() { - const info = getMcodeServerInfo(); - return (info && info.version) || "unknown"; -} - -export function handleHealth(_req, res) { +export async function handleHealth(_req, res) { + const { version } = await readEngineVersion(); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end( JSON.stringify({ @@ -32,7 +26,7 @@ export function handleHealth(_req, res) { defaultModel: DEFAULT_MODEL, defaultWorkspace: DEFAULT_WORKSPACE, mcodeCmd: MCODE_CMD, - mcodeVersion: engineVersion(), + mcodeVersion: version, maxConcurrent: MAX_CONCURRENT, }), ); diff --git a/packages/webui/server/routes/plugins.js b/packages/webui/server/routes/plugins.js index 1f24b6da..01023a34 100644 --- a/packages/webui/server/routes/plugins.js +++ b/packages/webui/server/routes/plugins.js @@ -19,8 +19,9 @@ // stays in the runtime. // // Data source: the catalogue host's `cliService`, reached through the -// exported `getCatalogueHost()` singleton in `lib/acp-client.js`. The host is -// booted unconditionally on first call, on purpose: +// engine facade's `getEngineCatalogueHost()` (server/engine/host.js), which +// forwards to the `getCatalogueHost()` singleton in `lib/acp-client.js`. +// The host is booted unconditionally on first call, on purpose: // // - `MCODE_WEBUI_TRANSPORT` defaults to `acp`, and `transportWantsCatalogue()` // only gates *session-list* traffic. ACP has no plugin method at all, so @@ -47,7 +48,7 @@ // ("official" | "local") so the webapp never has to import the protocol // package to tell the two apart (`@mavis/webui` does not depend on it). -import { getCatalogueHost } from "../lib/acp-client.js"; +import { getEngineCatalogueHost } from "../engine/index.js"; import { readJson } from "../lib/read-json.js"; /** Page size when the caller sends no `limit`; matches the facade default. */ @@ -78,7 +79,7 @@ function json(res, status, payload) { /** The default data source: the catalogue host singleton's bare cliService. */ async function defaultGetCliService() { - const host = await getCatalogueHost(); + const host = await getEngineCatalogueHost(); return host ? host.cliService : null; } diff --git a/packages/webui/server/routes/protocol.js b/packages/webui/server/routes/protocol.js index 4d0de38c..763a1036 100644 --- a/packages/webui/server/routes/protocol.js +++ b/packages/webui/server/routes/protocol.js @@ -18,9 +18,14 @@ import { cancelSession, loadSession, activateSession, - listSessions, mcodePermissionToWebui, } from "../lib/mcode-rpc.js"; +// M3-B1 (engine facade): only #72 (`list-sessions`) is gated in this +// batch. The other five handlers here still call mcode-rpc directly — +// they belong to B4 (#73 capabilities) and B7/B9 (cancel, load, activate, +// set-mode, set-config-option), each of which lands its own facade call +// with its own regression evidence. +import { readEngineSessionList } from "../engine/session-reads.js"; import { loadSessions, saveSessions, resetContext } from "../lib/sessions.js"; import { pushStateFor } from "../lib/state-bus.js"; import { readJson } from "../lib/read-json.js"; @@ -227,11 +232,21 @@ export async function handleActivateSession(req, res, ctx) { // ============================================================ // GET /api/protocol/list-sessions?cwd=... // 列 mcode session, 供前端 "远控 TUI" UI 用 +// +// M3-B1: the list now comes from the engine facade +// (`server/engine/session-reads.js`) instead of `mcode-rpc.js#listSessions` +// directly, so this endpoint is gated on the same declared +// `sessionCrud.listSessions` as the sidebar's #9 and #72 share. The +// facade forwards to the same `listAllMcodeSessions()` the rpc wrapper +// called, which means the runtime path already went through +// `lib/catalogue-sessions.js`; the cwd filter below and the response +// shape are untouched — `mcode-rpc.js#listSessions` is still exported +// for the write-family callers that arrive with later batches. // ============================================================ export async function handleListSessions(req, res, ctx) { const url = new URL(req.url, "http://localhost"); const cwd = url.searchParams.get("cwd") || ctx?.cs?.workspace?.dir || ""; - const all = await listSessions(); + const { sessions: all } = await readEngineSessionList(); if (!cwd) return respond(res, 200, { ok: true, sessions: all }); // 按 cwd 过滤 (norm 路径对齐) const norm = (p) => diff --git a/packages/webui/server/routes/sessions.js b/packages/webui/server/routes/sessions.js index 8638ffb4..b986f7e0 100644 --- a/packages/webui/server/routes/sessions.js +++ b/packages/webui/server/routes/sessions.js @@ -16,7 +16,6 @@ import { import { deleteMcodeSessionFromDb } from "../lib/mcode-session-delete.js"; import { getMcodeSessionTitle, - getMcodeSessionsForWorkspace, getMcodeSessionsCacheSync, getMcodeSessionsStaleSync, shutdownMcodeAcpSingleton, @@ -34,7 +33,30 @@ import { runChatViewChat, } from "../lib/state-bus.js"; import { MCODE_RUNTIME_DB, DEFAULT_WORKSPACE } from "../lib/config.js"; -import { getSessionTree, invalidateSessionTree } from "../lib/session-tree.js"; +import { invalidateSessionTree } from "../lib/session-tree.js"; +// M3-B1 (engine facade): #9 and #10 read the engine through the declared +// capability rather than straight off the ACP client. Both facade +// functions forward to the same acp-client exports this module already +// imported, so the wire shape, the cache and the transport switch are +// unchanged — only the gate in front of them is new. +import { + readEngineSessionListForWorkspace, + readEngineSessionTitle, +} from "../engine/session-reads.js"; +// M3-B2 (engine facade): #8 asks the facade, which checks the provider's +// declaration (sessionCrud.listSessions → 501 when absent) and then +// forwards to the same `getSessionTree` this module used to call +// directly. `invalidateSessionTree` stays a direct import: it is a +// synchronous cache drop with no I/O, it is called from the rename and +// delete paths, and routing a one-line invalidation through an async +// facade would make those paths wait on a module load to do nothing. +import { readEngineSessionTree } from "../engine/session-tree-reads.js"; +// The capability-error predicate `handleSessionTree` uses to tell the gate's +// 501 apart from a soft-fail. Taken from the facade entry, which re-exports +// the same binding `app.js#invokeHandler` matches on, so the two ends of this +// protocol cannot drift onto two different notions of "is this the gate's +// error". +import { isEngineCapabilityNotSupportedError } from "../engine/index.js"; import { authorize } from "../lib/authorize.js"; import { pushAlert } from "../lib/alerts.js"; import { append as _eventsAppend } from "../lib/events.js"; @@ -1011,13 +1033,34 @@ export async function handleDeleteSession(req, res, ctx) { // `?refresh=1` bypasses the 15s cache. A db that cannot be read is not a client // error: `ok:false` + `reason` lets the sidebar fall back to the wrapper list // instead of rendering an empty tree. -export function handleSessionTree(req, res, _ctx) { +// +// M3-B2: the read goes through the engine facade, which gates it on the +// provider's declared `sessionCrud.listSessions` and then forwards to the very +// same `getSessionTree`. The payload below is `tree` verbatim — same keys, same +// node shape, same `ok:false` soft-fail. The subtree hierarchy is built by +// `buildTree` from `parent_session_id` and is NOT re-derived here; a child that +// fails to attach to its parent is a subagent the user cannot see, so the tree +// has exactly one assembler and it is not this route. +export async function handleSessionTree(req, res, _ctx) { const url = new URL(req.url, "http://localhost"); const force = url.searchParams.get("refresh") === "1"; let payload; try { - payload = getSessionTree({ force }); + ({ tree: payload } = await readEngineSessionTree({ force })); } catch (cause) { + // Re-throw the capability gate, and only it. `invokeHandler` maps + // `EngineCapabilityNotSupportedError` to 501 — the deliberate "this + // provider cannot list sessions" answer — whereas this catch exists + // for the OTHER failures (a db that cannot be read, an assembler bug), + // which the sidebar is built to degrade on. Folding the capability + // error in here would answer `200 {ok:false}` to a request the server + // is refusing on purpose: the fake success the gate exists to prevent. + // + // The test is the class's own `instanceof` helper, not a `.name` + // compare. `name` is a writable instance property, so one stray + // `err.name = "…"` upstream would silently turn that 501 back into the + // soft failure — a failure mode that reads as a passing test. + if (isEngineCapabilityNotSupportedError(cause)) throw cause; payload = { ok: false, reason: "session_tree_failed", @@ -1036,17 +1079,32 @@ export function handleSessionTree(req, res, _ctx) { } // GET /api/acp-sessions?cwd=... — mcode acp session/list +// +// M3-B1: the read goes through the engine facade +// (engine/session-reads.js) so the sidebar's data source is a DECLARED +// capability rather than "whatever the transport happens to be". A +// provider that does not declare `sessionCrud.listSessions` answers 501 +// through app.js#invokeHandler instead of an empty list. The response +// shape is byte-for-byte what it was: the facade forwards to the same +// `getMcodeSessionsForWorkspace` (same 30s cache, same cwd +// normalisation, same `catalogue-sessions.js` projection on the runtime +// path). export async function handleAcpSessions(req, res, ctx) { const cs = ctx.cs; const url = new URL(req.url, "http://localhost"); const cwd = url.searchParams.get("cwd") || (cs.workspace && cs.workspace.dir) || ""; - const sessions = await getMcodeSessionsForWorkspace(cwd); + const { sessions } = await readEngineSessionListForWorkspace({ cwd }); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end(JSON.stringify({ ok: true, cwd, sessions })); } // GET /api/acp-session-title?sessionId=... +// +// M3-B1: gated on `sessionCrud.getSession` — the engine method the ACP +// `session/list` title lookup corresponds to. `title` stays `null` for +// both "no such session" and "engine has no title": the endpoint has +// always collapsed those two and callers depend on it. export async function handleAcpSessionTitle(req, res, _ctx) { const url = new URL(req.url, "http://localhost"); const sid = url.searchParams.get("sessionId") || ""; @@ -1054,7 +1112,7 @@ export async function handleAcpSessionTitle(req, res, _ctx) { res.writeHead(400, { "Content-Type": "application/json" }); return res.end(JSON.stringify({ ok: false, error: "sessionId required" })); } - const title = await getMcodeSessionTitle(sid); + const { title } = await readEngineSessionTitle({ sessionId: sid }); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end( JSON.stringify({ ok: true, sessionId: sid, title: title || null }), diff --git a/packages/webui/server/routes/state.js b/packages/webui/server/routes/state.js index aa6cc17f..796625c3 100644 --- a/packages/webui/server/routes/state.js +++ b/packages/webui/server/routes/state.js @@ -14,10 +14,13 @@ import { sessionsListForSnapshot, nextRevisionFor, } from "../lib/state-bus.js"; -import { - getMcodeSessionsForWorkspace, - getCachedMcodeCommands, -} from "../lib/acp-client.js"; +import { getCachedMcodeCommands } from "../lib/acp-client.js"; +// M3-B1 (engine facade): the declared-capability gate in front of the +// mcodeSessions mirror. The SSE first frame below keeps calling +// `mcodeSessionsSnapshotFields` directly — the SSE channel is a P2 +// migration, out of scope for this batch, and it must keep its exact +// pending/stale semantics. +import { readEngineSessionListForWorkspace } from "../engine/session-reads.js"; import { getLanBroadcast } from "../lib/settings.js"; import { applyMavisUsageToCs } from "../lib/mavis-usage.js"; import { getMcodeModelLimit } from "../lib/models.js"; @@ -110,9 +113,20 @@ const SSE_HEADERS = { export async function handleState(req, res, ctx) { const cs = getClient(ctx.cid); - const mcodeSessions = await getMcodeSessionsForWorkspace( - cs.workspace && cs.workspace.dir, - ); + // M3-B1: the mcodeSessions mirror now comes from the engine facade, + // which gates it on the declared `sessionCrud.listSessions` and reports + // (in the return value, not on the wire) whether the in-process host or + // the ACP mirror answered. The VALUE is the same array the endpoint + // built before — `readEngineSessionListForWorkspace` forwards to the + // same `getMcodeSessionsForWorkspace`, cache and cwd normalisation + // included. The snapshot body below is unchanged field for field: + // `snapshotViewFields` / `mcodeSessionsSnapshotFields` are the + // frontend's first-frame contract and this batch adds and removes + // nothing. + const { sessions: mcodeSessions } = await readEngineSessionListForWorkspace({ + cwd: (cs.workspace && cs.workspace.dir) || "", + endpoint: "GET /api/state", + }); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); // v0.5.bx-29: /api/state 也尝试 hydrate mavis db 真值 (best-effort) // SSE 客户端 (EventSource) 也会调这个端点, 所以 hydrate 也能发生在 reconnect 时 diff --git a/packages/webui/server/routes/turn-diff.js b/packages/webui/server/routes/turn-diff.js index bc2ee628..e5cbbbcc 100644 --- a/packages/webui/server/routes/turn-diff.js +++ b/packages/webui/server/routes/turn-diff.js @@ -8,8 +8,9 @@ // Zero new backend. Every endpoint is a thin projection over // `applications.session.diff` (getTurnDiff / revertTurnDiff / reapplyTurnDiff) // on the catalogue host — the same runtime application `routes/plugins.js` -// reaches through `getCatalogueHost()`. This file owns input validation, the -// wire shape, and the post-mutation refresh; it owns no diff logic. +// reaches through the engine facade's `getEngineCatalogueHost()`. This file +// owns input validation, the wire shape, and the post-mutation refresh; it +// owns no diff logic. // // Two deliberate constraints, both from the real-run verification in // `.tickets/webui-parity/82-coord-premise-verification.md`: @@ -43,7 +44,7 @@ // "Only the latest turn diff can be changed" / content-conflict gate, and the // card shows that message verbatim instead of a generic failure. -import { getCatalogueHost } from "../lib/acp-client.js"; +import { getEngineCatalogueHost } from "../engine/index.js"; import { readJson } from "../lib/read-json.js"; import { invalidateSessionTree } from "../lib/session-tree.js"; import { @@ -96,7 +97,7 @@ function readSelector(source) { /** `applications.session.diff` — and only that. */ async function defaultGetDiffApplication() { - const host = await getCatalogueHost(); + const host = await getEngineCatalogueHost(); const diff = host && host.applications ? host.applications.session?.diff : undefined; return diff ?? null; } diff --git a/packages/webui/server/routes/usage.js b/packages/webui/server/routes/usage.js index 1e8aae03..b4f16b62 100644 --- a/packages/webui/server/routes/usage.js +++ b/packages/webui/server/routes/usage.js @@ -12,26 +12,28 @@ // now decided by lib/usage.js#runUsageQuery's `record` option, so a // caller that is only rendering the number does not add a sample. Also // added handleForecast which exposes the prediction to the UI. +// +// M3-B3: all four usage endpoints now reach the engine through +// `engine/usage-reads.js` instead of naming lib/usage.js, lib/mavis-usage.js, +// lib/quota-forecast.js and lib/config.js themselves. Nothing about the +// wire changed — the facade forwards the payloads and owns the two +// DERIVED figures (`contextUsed`, the forecast) so the formulas have one +// home. See engine/usage-reads.js for why #19 declares no capability and +// why #17 does not yet call the provider's `getSessionUsage` method. -import { existsSync } from "node:fs"; -import { runUsageQuery } from "../lib/usage.js"; import { - getMavisTokenUsage, - getMavisTokenUsageModel, -} from "../lib/mavis-usage.js"; + readEngineAccountQuota, + readEngineQuotaForecast, + readEngineSessionUsage, +} from "../engine/usage-reads.js"; import { pushStateFor } from "../lib/state-bus.js"; import { getMcodeModelLimit } from "../lib/models.js"; -import { MAVIS_DB_PATH } from "../lib/config.js"; -// C07: quota exhaustion forecast (linear LS on usage history) -// readHistory + forecastExhaustion + recordSnapshotFromCs. -// Pure module — no state-bus / settings coupling, just FS + math. -import { readHistory, forecastExhaustion } from "../lib/quota-forecast.js"; import { readJson } from "../lib/read-json.js"; // POST /api/usage & /api/usage-trigger // -// The answer is the quota figures runUsageQuery just fetched. It used to be a +// The answer is the quota figures the read just fetched. It used to be a // bare {ok:true} written before the fetch — the popover reads this response // body, so it never saw a `remaining` even when the fetch succeeded. export async function handleUsage(req, res, ctx) { @@ -39,7 +41,9 @@ export async function handleUsage(req, res, ctx) { // history; the client's poll uses it. Absent or true means the historical // behaviour, where a read is also a measurement. const body = await readJson(req); - const payload = await runUsageQuery(ctx.cs, ctx.cid, { + const { payload } = await readEngineAccountQuota({ + cs: ctx.cs, + cid: ctx.cid, record: body.record !== false, }); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); @@ -68,26 +72,29 @@ export async function handleUsageReal(req, res, ctx) { }), ); } - const usage = await getMavisTokenUsage(sid); - const model = await getMavisTokenUsageModel(sid); - if (!usage) { + const read = await readEngineSessionUsage({ mcodeSessionId: sid }); + if (!read.found) { res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end( JSON.stringify({ ok: true, found: false, sid, - dbPath: MAVIS_DB_PATH, - dbExists: existsSync(MAVIS_DB_PATH), + dbPath: read.dbPath, + dbExists: read.dbExists, }), ); } + const usage = read.usage; res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end( JSON.stringify({ ok: true, found: true, sid, + // `rows` is a COUNT, not a list — the engine's provider method + // answers a row ARRAY under the same name, which is one of the + // reasons #17 does not call it yet (engine/usage-reads.js header). rows: usage.rows, totalInput: usage.totalInput, totalOutput: usage.totalOutput, @@ -95,12 +102,15 @@ export async function handleUsageReal(req, res, ctx) { totalCacheWrite: usage.totalCacheWrite, totalReasoning: usage.totalReasoning, // v0.5.bx-10 fix: context 实际是 input + output + reasoning (cache 是 input 子集) - contextUsed: usage.totalInput + usage.totalOutput + usage.totalReasoning, - model: (model && model.model) || null, + // CUMULATIVE, deliberately not the chat flow's per-turn + // `lastTurnContextTokens`. The formula now lives in the engine layer + // as `contextUsedTokens` and is pinned on its inputs there. + contextUsed: read.contextUsed, + model: read.model, modelLimit: getMcodeModelLimit(cs.model && cs.model.name), firstTs: usage.firstTs, lastTs: usage.lastTs, - dbPath: MAVIS_DB_PATH, + dbPath: read.dbPath, }), ); } @@ -108,20 +118,11 @@ export async function handleUsageReal(req, res, ctx) { // C07: GET /api/usage/forecast — predict quota exhaustion time. // Reads ~/.mcode-webui/usage-history.ndjson, runs forecastExhaustion, // and returns the JSON payload documented in CAPABILITIES.md §8. -// Best-effort: if the file is missing or empty, returns -// { ok: true, forecast: { ... reason: "no_history" } } so the UI -// can render a "collecting data…" placeholder instead of erroring. +// Best-effort: if the file is missing or empty, the read answers +// { … reason: "no_history" } so the UI can render a "collecting data…" +// placeholder instead of erroring. export async function handleForecast(_req, res, _ctx) { - let history = []; - try { - history = readHistory(); - } catch { - // readHistory already swallows FS errors; this catch is just a - // belt-and-braces guard so a buggy extension never breaks the - // endpoint. - history = []; - } - const forecast = forecastExhaustion(history); + const { forecast } = await readEngineQuotaForecast(); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end( JSON.stringify({ diff --git a/packages/webui/test/helpers/_setup.js b/packages/webui/test/helpers/_setup.js index 9ee72588..d30db95f 100644 --- a/packages/webui/test/helpers/_setup.js +++ b/packages/webui/test/helpers/_setup.js @@ -51,6 +51,12 @@ const _acpMock = { getMcodeAcpClient: async () => null, listAllMcodeSessions: async () => [], getMcodeServerInfo: () => null, + // M3-B1: the engine facade (server/engine/session-reads.js) asks whether + // the in-process catalogue host answered before it reports where a read's + // bytes came from. Default null = "the host never booted", i.e. the + // acp-fallback case. Tests that want the catalogue case register + // `getCatalogueHost: async () => ({ adapter: {} })`. + getCatalogueHost: async () => null, invalidateMcodeSessionsCache: () => {}, shutdownMcodeAcpSingleton: () => {}, dropMcodeSessionFromCache: () => {}, // v1.0: 删除路由防复活用 @@ -217,6 +223,9 @@ export async function setupMocks(t, overrides = {}) { getMcodeAcpClient: (...a) => _acpMock.getMcodeAcpClient(...a), listAllMcodeSessions: (...a) => _acpMock.listAllMcodeSessions(...a), getMcodeServerInfo: (...a) => _acpMock.getMcodeServerInfo(...a), + // M3-B1: engine/session-reads.js asks this to report whether a read + // came from the in-process catalogue host or from the ACP mirror. + getCatalogueHost: (...a) => _acpMock.getCatalogueHost(...a), invalidateMcodeSessionsCache: (...a) => _acpMock.invalidateMcodeSessionsCache(...a), shutdownMcodeAcpSingleton: (...a) => diff --git a/packages/webui/test/lib/engine/capability-snapshot.test.js b/packages/webui/test/lib/engine/capability-snapshot.test.js new file mode 100644 index 00000000..a2e66a53 --- /dev/null +++ b/packages/webui/test/lib/engine/capability-snapshot.test.js @@ -0,0 +1,492 @@ +// webui/test/lib/engine/capability-snapshot.test.js +// +// M2 — capability-declaration snapshot audit against the REAL host +// (design doc §2.4, migration step M2; doc/engine-abstraction-design.md). +// +// M1 (test/lib/engine/capabilities.test.js) pins every declared LEVEL +// against the audited matrix. That alone cannot catch the more dangerous +// drift: the declaration saying "full"/"partial" while the live object no +// longer carries the promised methods (or has grown the ones `missing` +// denies). This file closes that gap by booting ONE real catalogue host +// against an isolated tmp data dir and auditing every full/partial key +// against the reflected method surfaces: +// +// full → every REQUIRED_METHODS entry must be typeof "function" on +// the declared surface member; +// partial → methods of the key that ARE named in `missing` must be +// absent; the ones NOT named must be present; kebab-case +// `missing` items (sub-capability names such as "file-write") +// must have NO method on the surface whose name contains all +// their segments (a future getWorkspaceGitDiff would make the +// "git-diff" entry go red until the declaration is re-audited); +// none → not method-checked (a provider may legitimately expose no +// surface for the capability). +// +// The audit function is a PURE function over (declaration, method-name +// sets), so the mutation checks below feed it hand-built mutant surfaces +// and assert it reports the drift — the "flip a level / delete a method +// must go red" requirement is thereby pinned as a test of the checker +// itself, not just performed once by hand. +// +// Isolation: the host boots against a per-run tmp dir via mkTmpDir and +// MINIMAX_DATA_DIR / MCODE_WEBUI_* are pinned BEFORE the dynamic import +// of the engine provider (node:test runs each file in its own process; +// setting only MCODE_WEBUI_DATA_DIR is NOT enough — the engine dir would +// fall back to ~/.minimax and rewrite the user's real config). + +import { test, describe, before, after } from "node:test"; +import { strict as assert } from "node:assert"; + +import { mkTmpDir, rmTmpDir } from "../../helpers/tmp.js"; + +// Set BEFORE any dynamic import of config-reading / host modules below. +const tmpBase = mkTmpDir("mcode-webui-engine-snapshot-"); +process.env.MINIMAX_DATA_DIR = tmpBase; +process.env.MCODE_WEBUI_DATA_DIR = tmpBase; +process.env.MCODE_WEBUI_SETTINGS_PATH = `${tmpBase}/settings.json`; +process.env.MCODE_WEBUI_EVENTS_PATH = `${tmpBase}/events.jsonl`; +process.env.MCODE_WEBUI_SESSIONS_DB = `${tmpBase}/sessions.db`; +process.env.MCODE_WEBUI_UPLOAD_DIR = `${tmpBase}/uploads`; + +// Declaration modules are import-light (no @mavis/* tree), and the env +// above is already pinned, so loading them at top level is safe here. +const { + ENGINE_CAPABILITY_KEYS, + LOCAL_RUNTIME_V2_CAPABILITIES, + TUI_RUNTIME_ADAPTER_CAPABILITIES, + getEngineProvider, + listEngineProviderIds, + validateEngineCapabilities, +} = await import("../../../server/engine/index.js"); + +// --------------------------------------------------------------------------- +// REQUIRED_METHODS — what each capability key means ON THE OBJECTS. +// --------------------------------------------------------------------------- +// +// Provenance (how this table was derived, per ticket 104): a one-off +// audit script booted the real catalogue host exactly like this file +// does, walked the prototype chains of host.adapter / host.cliService / +// host.applications.session.diff with getOwnPropertyNames, and dumped +// the full method sets — 91 adapter methods, 94 CliService methods, and +// the session.diff facade (getSessionDiff/getTurnDiff/revertTurnDiff/ +// reapplyTurnDiff + the internal requireTarget). The lists below name +// exactly the methods each declaration's own evidence comments cite +// (server/engine/providers/*.js), each re-verified present/absent on +// those dumped sets. `on` is the host member the method must live on: +// the tui-runtime-adapter provider declares the adapter surface; the +// local-runtime-v2 provider declares cliService + applications. + +/** Which host member each provider's surface lives on. */ +const SURFACE_MEMBERS = { + "tui-runtime-adapter": ["adapter"], + "local-runtime-v2": ["cliService", "applications.session.diff"], +}; + +function resolveMember(host, dottedPath) { + return dottedPath.split(".").reduce((obj, key) => (obj == null ? obj : obj[key]), host); +} + +/** + * REQUIRED_METHODS[providerId][capabilityKey] pins what the capability + * key MEANS on that provider's surface: + * - `on`: the host member the key's methods live on; + * - `methods`: methods that MUST exist when the key is full (and, for + * a partial, the parts that are present); + * - `absent`: method-NAMED sub-items the partial declarations list in + * `missing` — methods of this capability's domain that genuinely do + * not exist on this surface (reapplyTurnDiff on the adapter, + * getDelegationSnapshot on the bare CliService). They are part of + * the snapshot so "missing must really be absent" is checked, and a + * partial that stops listing one goes red (under-declaration). + */ +const REQUIRED_METHODS = { + "tui-runtime-adapter": { + sessionCrud: { on: "adapter", methods: ["createSession", "listSessions", "getSession", "renameSession", "archiveSession", "deleteSession", "forkSession"] }, + streamingSend: { on: "adapter", methods: ["sendMessage", "watchSessionTurn", "watchEvents"] }, + interrupt: { on: "adapter", methods: ["abortSession", "steer"] }, + toolSkillInvocation: { on: "adapter", methods: ["listSkills", "listPendingPermissions", "replyPermission"] }, + turnRewindRedo: { on: "adapter", methods: ["rewindSession", "getSessionRewindPreview"], absent: ["reapplyTurnDiff"] }, + plugins: { on: "adapter", methods: ["listInstalledPlugins", "listMarketplacePlugins", "mutatePlugin", "refreshPlugins"], absent: ["previewGithubPlugin", "importGithubPlugin", "listEnabledPlugins"] }, + mcp: { on: "adapter", methods: ["configureSessionMcpServers", "clearSessionMcpServers", "inspectProjectMcp", "listMcpServers"] }, + subagents: { on: "adapter", methods: ["getDelegationSnapshot", "stopDelegation", "listBackgroundTasks"] }, + usageStats: { on: "adapter", methods: ["getSessionUsage", "getSessionUsageSummary", "watchSessionUsageCommits"] }, + authCredentials: { on: "adapter", methods: ["getAccountStatus", "getCodexOAuthStatus", "startCodexOAuthLogin", "cancelCodexOAuthLogin", "getMiniMaxApiKeyStatus", "upsertMiniMaxApiKey", "listUserModelProviders", "createUserModelProvider", "updateUserModelProvider", "deleteUserModelProvider", "testUserModelProvider", "discoverUserModelsCandidate"] }, + fileReadWrite: { on: "adapter", methods: ["listWorkspaceFileTree", "searchWorkspaceFiles"] }, + gitOperations: { on: "adapter", methods: ["getWorkspaceGitMetadata"] }, + }, + "local-runtime-v2": { + sessionCrud: { on: "cliService", methods: ["createSession", "updateSession", "archiveSession", "deleteSession", "forkSession", "getSessionForkOptions"] }, + streamingSend: { on: "cliService", methods: ["sendMessage", "resumeSession", "steerSession", "watchEvents"] }, + interrupt: { on: "cliService", methods: ["abortSession"] }, + toolSkillInvocation: { on: "cliService", methods: ["listSkills", "listRuntimeSkills", "listPendingPermissions", "replyPermission"] }, + turnDiff: { on: "applications.session.diff", methods: ["getSessionDiff", "getTurnDiff", "revertTurnDiff", "reapplyTurnDiff"] }, + turnRewindRedo: { on: "cliService", methods: ["getSessionRewindPreview", "rewindSession", "editSessionMessage"] }, + plugins: { on: "cliService", methods: ["refreshPlugins", "listMarketplacePlugins", "listInstalledPlugins", "listEnabledPlugins", "installPlugin", "enablePlugin", "disablePlugin", "uninstallPlugin", "previewGithubPlugin", "importGithubPlugin"] }, + mcp: { on: "cliService", methods: ["configureSessionMcpServers", "inspectProjectMcp", "clearSessionMcpServers", "listMcpServers"] }, + subagents: { on: "cliService", methods: ["listBackgroundTasks"], absent: ["getDelegationSnapshot", "stopDelegation"] }, + usageStats: { on: "cliService", methods: ["getSessionUsage", "getSessionUsageSummary", "watchSessionUsageCommits"] }, + authCredentials: { on: "cliService", methods: ["getAccountStatus", "getCodexOAuthStatus", "startCodexOAuthLogin", "cancelCodexOAuthLogin", "getMiniMaxApiKeyStatus", "upsertMiniMaxApiKey", "listUserModelProviders", "createUserModelProvider", "updateUserModelProvider", "deleteUserModelProvider", "testUserModel", "discoverUserModelsCandidate"] }, + fileReadWrite: { on: "cliService", methods: ["listWorkspaceFileTree", "searchWorkspaceFiles"] }, + gitOperations: { on: "cliService", methods: ["getWorkspaceGitMetadata", "getWorkspaceReviewLink"] }, + }, +}; + +// --------------------------------------------------------------------------- +// The pure audit — errors are values (a problems list), so the mutation +// checks can feed it synthetic surfaces and pin that it reports drift. +// --------------------------------------------------------------------------- + +/** + * Walk an object's prototype chain and collect every own function name + * (skipping Object.prototype noise). This is the same reflection the +// one-off provenance audit used, so "exists" means exactly what the + * table was derived against — class methods live on prototypes, so a + * plain Object.keys() would see none of them. + */ +export function collectMethodNames(obj) { + const names = new Set(); + let proto = obj; + const seen = new Set(); + while (proto && proto !== Object.prototype && !seen.has(proto)) { + seen.add(proto); + for (const name of Object.getOwnPropertyNames(proto)) { + if (name === "constructor") continue; + try { + if (typeof obj[name] === "function") names.add(name); + } catch { + // getter that throws — not a method + } + } + proto = Object.getPrototypeOf(proto); + } + return [...names].sort(); +} + +/** + * Does any method name on the surface cover all segments of a + * kebab-case sub-capability name ("file-write" → ["file","write"])? + * Both segments must appear in the SAME method name: getWorkspaceGit- + * Metadata contains "git" but not "diff", so it does not satisfy + * "git-diff"; a future getWorkspaceGitDiff would. + */ +function subCapabilityHasMethods(missingItem, allMethodNames) { + const segments = missingItem.split("-").map((s) => s.toLowerCase()); + return allMethodNames.filter((name) => { + const lower = name.toLowerCase(); + return segments.every((segment) => lower.includes(segment)); + }); +} + +/** + * Audit one provider's declaration against the live host. + * + * @param {string} providerId + * @param {Record} declaration + * @param {object} host the real catalogue host (adapter/cliService/…) + * @returns {string[]} problems; empty means the declaration matches the + * implementation for every full/partial key. + */ +export function auditProviderCapabilities(providerId, declaration, host) { + const problems = []; + const required = REQUIRED_METHODS[providerId] || {}; + const surfaceMethodsByMember = new Map(); + const methodTypeOf = (on, method) => { + const member = resolveMember(host, on); + if (member === undefined || member === null) return "undefined"; + try { + return typeof member[method]; + } catch { + return "throws"; + } + }; + const surfaceMethodNames = (on) => { + if (!surfaceMethodsByMember.has(on)) { + const member = resolveMember(host, on); + surfaceMethodsByMember.set(on, member ? collectMethodNames(member) : []); + } + return surfaceMethodsByMember.get(on); + }; + + for (const key of Object.keys(required)) { + const entry = declaration[key]; + if (!entry) continue; // shape problems are M1's validate, not this audit + const { on, methods, absent = [] } = required[key]; + + if (entry.level === "full") { + for (const method of methods) { + if (methodTypeOf(on, method) !== "function") { + problems.push( + `${providerId}.${key}: declared full but ${on}.${method} is not a function`, + ); + } + } + continue; + } + + if (entry.level === "partial") { + const missing = entry.missing || []; + // Present part: every tracked method must exist (none of them may + // appear in `missing` — see the coverage sweep below). + for (const method of methods) { + if (methodTypeOf(on, method) !== "function") { + problems.push( + `${providerId}.${key}: declared partial, not listing ${on}.${method} as missing, yet it is absent`, + ); + } + } + // Absent part: each method-named missing item must be tracked + // (else the audit would be vacuous for it) and genuinely absent. + for (const item of missing) { + if (item.includes("-")) continue; // sub-capability name, swept below + if (!absent.includes(item)) { + problems.push( + `${providerId}.${key}: missing lists "${item}" which this snapshot does not track as absent for the key`, + ); + continue; + } + if (methodTypeOf(on, item) === "function") { + problems.push( + `${providerId}.${key}: missing lists ${on}.${item} but it exists on the surface`, + ); + } + } + // Under-declaration: a tracked absent method the declaration + // stopped listing would hide a real gap behind "partial". + for (const item of absent) { + if (!missing.includes(item)) { + problems.push( + `${providerId}.${key}: ${on}.${item} is absent from the surface but the declaration does not list it as missing`, + ); + } + } + // Kebab-case missing items name sub-capabilities, not methods: + // they must have NO covering method on the key's surface. + for (const item of missing) { + if (!item.includes("-")) continue; + const covered = subCapabilityHasMethods(item, surfaceMethodNames(on)); + if (covered.length > 0) { + problems.push( + `${providerId}.${key}: missing lists sub-capability "${item}" but surface method(s) ${covered.join(", ")} cover it`, + ); + } + } + continue; + } + // "none": deliberately not method-checked. + } + return problems; +} + +// --------------------------------------------------------------------------- +// Static guard — registry-driven key-set assertion (no host needed). +// --------------------------------------------------------------------------- + +describe("M2 static guard — declarations carry exactly the 14 contract keys", () => { + test("every REGISTERED provider declares exactly ENGINE_CAPABILITY_KEYS — no typos can pass silently", () => { + // Registry-driven on purpose: M4 will register acp/exec providers, + // and this sweep picks them up without editing the test. A key the + // contract does not know (typo, rename) or a dropped key fails here + // even before any host is booted. + const ids = listEngineProviderIds(); + assert.ok(ids.length >= 2, `expected both M1 providers registered, got ${ids.join(", ")}`); + const expected = [...ENGINE_CAPABILITY_KEYS].sort(); + for (const id of ids) { + const { capabilities } = getEngineProvider(id); + assert.deepEqual( + Object.keys(capabilities).sort(), + expected, + `${id} must declare exactly the 14 contract keys`, + ); + assert.deepEqual( + validateEngineCapabilities(capabilities), + [], + `${id} declaration must pass contract validation`, + ); + } + }); + + test("REQUIRED_METHODS covers every non-none key of every audited provider (and no others)", () => { + for (const [providerId, required] of Object.entries(REQUIRED_METHODS)) { + const { capabilities } = getEngineProvider(providerId); + for (const key of Object.keys(required)) { + assert.ok( + capabilities[key] && capabilities[key].level !== "none", + `${providerId}.${key} is audited but declared none — none keys are not method-checked`, + ); + assert.ok( + SURFACE_MEMBERS[providerId].includes(required[key].on) || + required[key].on.startsWith("applications."), + `${providerId}.${key} surface "${required[key].on}" must be a declared surface member`, + ); + } + } + }); +}); + +// --------------------------------------------------------------------------- +// Live-host audit — one real catalogue host, both providers audited. +// --------------------------------------------------------------------------- + +describe("M2 snapshot — declarations vs the REAL catalogue host", () => { + let host; + let declarations; + + before(async () => { + // Dynamic import AFTER env is pinned: the provider module pulls the + // @mavis/* TS tree and constructs the real in-process runtime. + const { createCatalogueHost } = await import( + "../../../server/engine/providers/local-runtime-v2.js" + ); + declarations = { + "local-runtime-v2": LOCAL_RUNTIME_V2_CAPABILITIES, + "tui-runtime-adapter": TUI_RUNTIME_ADAPTER_CAPABILITIES, + }; + host = await createCatalogueHost({ dataDir: tmpBase }); + }); + + after(async () => { + if (host) await host.close(); + rmTmpDir(tmpBase); + }); + + test("the host exposes the surfaces the declarations talk about", () => { + // Precondition tripwire: if the host contract loses a member the + // audit below would silently degrade to checking nothing. + assert.equal(typeof host.adapter?.sendMessage, "function", "host.adapter missing"); + assert.equal(typeof host.cliService?.createSession, "function", "host.cliService missing"); + assert.equal( + typeof host.applications?.session?.diff?.getTurnDiff, + "function", + "host.applications.session.diff missing", + ); + }); + + for (const providerId of Object.keys(REQUIRED_METHODS)) { + test(`${providerId}: every full/partial key matches the live surface (none keys unchecked)`, () => { + const problems = auditProviderCapabilities( + providerId, + declarations[providerId], + host, + ); + assert.deepEqual( + problems, + [], + `declaration/implementation drift must be empty — a non-empty list is the CI red light M2 exists for:\n ${problems.join("\n ")}`, + ); + }); + } + + test("method-surface sizes stay in the audited ballpark (gross-loss tripwire)", () => { + // Not an exact pin (the engine may add methods freely) — this only + // catches a wholesale surface loss (e.g. a proxy/wrapper hiding the + // prototype chain) that per-method checks above could otherwise + // never distinguish from a legitimately smaller surface. + assert.ok(collectMethodNames(host.adapter).length > 80, "adapter surface collapsed"); + assert.ok(collectMethodNames(host.cliService).length > 85, "cliService surface collapsed"); + }); +}); + +// --------------------------------------------------------------------------- +// Mutation checks — the checker itself must go red on drift. These pin +// the ticket's mutation matrix against synthetic surfaces, so the red +// light is guaranteed by tests, not by a one-time manual run. +// --------------------------------------------------------------------------- + +describe("M2 mutation checks — auditProviderCapabilities reports drift", () => { + /** A minimal fake host from method-name lists per surface member. */ + function fakeHost(adapterNames, cliServiceNames, diffNames) { + const toObject = (names) => + Object.fromEntries(names.map((n) => [n, () => {}])); + return { + adapter: toObject(adapterNames), + cliService: toObject(cliServiceNames), + applications: { session: { diff: toObject(diffNames) } }, + }; + } + + const ADAPTER_ALL = REQUIRED_METHODS["tui-runtime-adapter"]; + const V2_ALL = REQUIRED_METHODS["local-runtime-v2"]; + + /** Method names per surface member, gathered from REQUIRED_METHODS. */ + function namesBySurface(provider) { + const byOn = {}; + for (const { on, methods } of Object.values(provider)) { + byOn[on] = [...(byOn[on] || []), ...methods]; + } + return byOn; + } + + test("MUT-1: flipping a full to partial (missing a method that EXISTS) goes red", () => { + // usageStats exists in full on cliService; declaring it partial and + // listing getSessionUsage as missing must fail the audit — this is + // the ticket's "flip a full to partial → red" mutation, pinned as a + // property of the checker. + const v2 = namesBySurface(V2_ALL); + const mutated = { + ...LOCAL_RUNTIME_V2_CAPABILITIES, + usageStats: { level: "partial", missing: ["getSessionUsage"], reason: "mutant" }, + }; + const problems = auditProviderCapabilities( + "local-runtime-v2", + mutated, + fakeHost([], v2.cliService, v2["applications.session.diff"]), + ); + assert.ok( + problems.some((p) => p.includes("usageStats") && p.includes("getSessionUsage")), + `expected the full→partial flip to be reported, got: ${JSON.stringify(problems)}`, + ); + }); + + test("MUT-2: deleting a method implementation goes red (full key)", () => { + const byOn = namesBySurface(V2_ALL); + const withoutDisablePlugin = byOn.cliService.filter((m) => m !== "disablePlugin"); + const problems = auditProviderCapabilities( + "local-runtime-v2", + LOCAL_RUNTIME_V2_CAPABILITIES, + fakeHost([], withoutDisablePlugin, byOn["applications.session.diff"]), + ); + assert.ok( + problems.some((p) => p.includes("plugins") && p.includes("disablePlugin")), + `expected the deleted method to be reported, got: ${JSON.stringify(problems)}`, + ); + }); + + test("MUT-3: deleting a method a partial relies on goes red", () => { + const byOn = namesBySurface(ADAPTER_ALL); + const withoutRewind = byOn.adapter.filter((m) => m !== "rewindSession"); + const problems = auditProviderCapabilities( + "tui-runtime-adapter", + TUI_RUNTIME_ADAPTER_CAPABILITIES, + fakeHost(withoutRewind, [], []), + ); + assert.ok( + problems.some((p) => p.includes("turnRewindRedo") && p.includes("rewindSession")), + `expected the deleted partial method to be reported, got: ${JSON.stringify(problems)}`, + ); + }); + + test("MUT-4: a missing sub-capability that GREW a covering method goes red", () => { + // The engine grows getWorkspaceGitDiff while the declaration still + // denies "git-diff" — the snapshot must force a re-audit. + const byOn = namesBySurface(ADAPTER_ALL); + const problems = auditProviderCapabilities( + "tui-runtime-adapter", + TUI_RUNTIME_ADAPTER_CAPABILITIES, + fakeHost([...byOn.adapter, "getWorkspaceGitDiff"], [], []), + ); + assert.ok( + problems.some((p) => p.includes("gitOperations") && p.includes("getWorkspaceGitDiff")), + `expected the grown sub-capability to be reported, got: ${JSON.stringify(problems)}`, + ); + }); + + test("MUT-5: a partial listing an absent method as missing is fine; listing a present one is not", () => { + const byOn = namesBySurface(ADAPTER_ALL); + const ok = auditProviderCapabilities( + "tui-runtime-adapter", + TUI_RUNTIME_ADAPTER_CAPABILITIES, + fakeHost(byOn.adapter, [], []), + ); + assert.deepEqual(ok, [], "the pristine declaration over the real method set is clean"); + }); +}); diff --git a/packages/webui/test/lib/engine/host-facade.test.js b/packages/webui/test/lib/engine/host-facade.test.js new file mode 100644 index 00000000..48cb1d64 --- /dev/null +++ b/packages/webui/test/lib/engine/host-facade.test.js @@ -0,0 +1,264 @@ +// webui/test/lib/engine/host-facade.test.js +// +// Regression guard for migration step M3, batch B0: the catalogue host is +// reached through the engine facade, and the facade itself stays on the +// right side of the boot-path boundary. +// +// The M1 lesson is why this file exists. Moving the plugins and turn-diff +// endpoints onto the facade looks like a rename, and the tempting way to +// write it is a static `import { getCatalogueHost } from +// "../lib/acp-client.js"` inside `engine/host.js`. That compiles and passes +// every handler test — they inject `deps.getCliService` / +// `deps.getDiffApplication`, so the default getter never runs — while +// putting the ACP client tree behind `engine/index.js`, which `app.js` loads +// at boot. M1 already paid for that mistake once (209ms → 2700ms; the +// facade's own load 4685ms → 5ms after declaration and construction were +// split into two files). +// +// So the assertions come in two kinds, and the second is the load-bearing +// one: +// +// 1. Source shape — the two route files name the facade and never +// `lib/acp-client.js`; `engine/host.js` reaches the singleton through a +// dynamic import; `engine/index.js` re-exports the getter. +// 2. The real module graph — a fresh child process installs a +// `module.registerHooks` resolve hook, imports one entry, and reports +// every specifier the loader was asked to resolve, per parent. That +// yields the entry's direct edges and its transitive closure without +// guessing from the source text. A timing assertion would pass on a +// fast machine and fail on a loaded one; the module graph is a fact. +// +// Scope note, so this file is not mistaken for a global invariant: +// `routes/turn-diff.js` still pulls `lib/acp-client.js` TRANSITIVELY, through +// `lib/state-bus.js` (which app.js loads anyway). B0 removes the two direct +// edges; closing the state-bus one belongs to the batches that route the +// catalogue read/write families (M3-B1+), not here. The turn-diff assertions +// are therefore about its direct edges only, and say so. + +import { test, describe } from "node:test"; +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import { execFileSync } from "node:child_process"; +import { join, relative } from "node:path"; +import { fileURLToPath, pathToFileURL } from "node:url"; + +const packageDir = join(import.meta.dirname, "..", "..", ".."); +const serverDir = join(packageDir, "server"); +const read = (rel) => readFileSync(join(serverDir, rel), "utf8"); + +/** Strip comments — the prose in these files legitimately names the modules under guard. */ +function codeOf(source) { + return source + .replace(/\/\*[\s\S]*?\*\//g, "") + .split("\n") + .map((line) => line.replace(/^\s*\/\/.*$/, "")) + .join("\n"); +} + +/** + * Product files that carry the runtime / ACP tree. Loading any of them from + * the facade, whatever the reason given, is the regression. + */ +const FORBIDDEN_ON_BOOT_PATH = new Set([ + "lib/acp-client.js", + "lib/runtime-host.js", + "acp.mjs", + "providers/local-runtime-v2.js", +]); + +/** Bare specifiers whose package alone is enough to blow the boot budget. */ +const HEAVY_PACKAGE_PREFIXES = ["@mavis/", "@minimax/"]; + +// --- 1. source shape ------------------------------------------------------- + +describe("the host-consuming routes reach the host through the facade", () => { + for (const route of ["routes/plugins.js", "routes/turn-diff.js"]) { + test(`${route} imports the facade, not the host singleton module`, () => { + const code = codeOf(read(route)); + assert.ok( + !/from\s+"\.\.\/lib\/acp-client\.js"/.test(code), + `${route} must not import lib/acp-client.js directly — take the host from ../engine/index.js`, + ); + assert.ok( + /import\s*\{[^}]*getEngineCatalogueHost[^}]*\}\s*from\s+"\.\.\/engine\/index\.js"/.test(code), + `${route} must import getEngineCatalogueHost from ../engine/index.js`, + ); + assert.match(code, /await getEngineCatalogueHost\(\)/); + }); + + test(`${route} never names getCatalogueHost`, () => { + // Belt and braces: a future re-export of the raw getter from the facade + // would satisfy the import assertion above while quietly restoring the + // old name. The call site is what has to move. + assert.ok( + !/\bgetCatalogueHost\b/.test(codeOf(read(route))), + `${route} must not mention getCatalogueHost — the facade getter is getEngineCatalogueHost`, + ); + }); + } + + test("engine/host.js reaches the singleton through a dynamic import only", () => { + const code = codeOf(read("engine/host.js")); + assert.ok( + /await\s+import\(\s*"\.\.\/lib\/acp-client\.js"\s*\)/.test(code), + "the facade must load lib/acp-client.js with await import()", + ); + // A static `import ... from "../lib/acp-client.js"` anywhere in this file + // — even one used for nothing but a type — puts the module back on the + // boot path, because app.js loads engine/index.js. + assert.ok( + !/^\s*import\s[^\n]*"\.\.\/lib\/acp-client\.js"/m.test(code), + "engine/host.js must not statically import lib/acp-client.js", + ); + }); + + test("engine/index.js re-exports the facade getter", () => { + // Unexported, both routes would import `undefined` and throw on the first + // real request — a failure no handler test reaches, since they inject + // their own data source. + assert.match( + read("engine/index.js"), + /export\s*\{[^}]*getEngineCatalogueHost[^}]*\}\s*from\s*"\.\/host\.js"/, + "engine/index.js must re-export getEngineCatalogueHost from ./host.js", + ); + }); +}); + +// --- 2. the real module graph ---------------------------------------------- + +/** + * Import `entryRel` in a fresh Node process and report its module graph: + * the specifiers resolved with the entry as their direct parent, the + * transitive set of product files, and the bare package specifiers. + * + * `module.registerHooks` is in-thread and unflagged on every Node this + * package supports (engines: >=22.19), so this needs no loader file and no + * experimental flag. The child mirrors the server's own source-layout + * bootstrap (`registerWorkspaceSources`, see server/lib/workspace-sources.js) + * and installs the hook AFTER it, so the recorded set is the entry's graph + * and not the harness's. The payload is framed by a sentinel because + * importing the graph legitimately prints lines of its own (lib/config.js + * logs the resolved workspace on load). + */ +function moduleGraphOf(entryRel) { + const entryUrl = pathToFileURL(join(serverDir, entryRel)).href; + const sentinel = "__MODULE_GRAPH__"; + const script = ` + import { registerHooks } from "node:module"; + const { registerWorkspaceSources } = await import(${JSON.stringify( + pathToFileURL(join(serverDir, "lib/workspace-sources.js")).href, + )}); + registerWorkspaceSources(); + const seen = []; + registerHooks({ + resolve(specifier, context, nextResolve) { + const resolved = nextResolve(specifier, context); + seen.push({ specifier, url: resolved.url, parent: context.parentURL }); + return resolved; + }, + }); + const entry = await import(${JSON.stringify(entryUrl)}); + process.stdout.write("\\n${sentinel}" + JSON.stringify({ seen, exports: Object.keys(entry) })); + `; + const stdout = execFileSync(process.execPath, ["--input-type=module", "-e", script], { + cwd: packageDir, + encoding: "utf8", + stdio: ["ignore", "pipe", "pipe"], + }); + const framed = stdout.slice(stdout.lastIndexOf(sentinel) + sentinel.length); + const { seen, exports: exportNames } = JSON.parse(framed); + + const isProductFile = (url) => url.startsWith(pathToFileURL(join(serverDir, "")).href); + const asProductPath = (url) => relative(serverDir, fileURLToPath(url)).split("\\").join("/"); + + return { + exportNames, + // Direct edges of the entry: what this file itself asks the loader for. + directSpecifiers: seen.filter((e) => e.parent === entryUrl).map((e) => e.specifier), + // Everything reachable, transitively. Product files only: a dependency's + // own internals are not what this guard is about. + productFiles: [...new Set(seen.filter((e) => isProductFile(e.url)).map((e) => asProductPath(e.url)))], + bareSpecifiers: [ + ...new Set( + seen + .map((e) => e.specifier) + .filter((s) => !s.startsWith(".") && !s.startsWith("file:") && !s.startsWith("node:")), + ), + ], + }; +} + +describe("the two routes resolve the host module through the facade only", () => { + for (const route of ["routes/plugins.js", "routes/turn-diff.js"]) { + test(`${route} has no direct edge to the acp client module`, () => { + // The resolved graph, not the source text: a barrel re-export that + // pulls acp-client in behind the facade would still show up here. + const graph = moduleGraphOf(route); + assert.ok( + graph.directSpecifiers.includes("../engine/index.js"), + `${route} must resolve ../engine/index.js directly (got: ${graph.directSpecifiers.join(", ")})`, + ); + for (const specifier of graph.directSpecifiers) { + assert.ok( + !/acp-client|runtime-host|acp\.mjs/.test(specifier), + `${route} must not resolve ${specifier} directly`, + ); + } + }); + } +}); + +describe("the facade stays off the heavy side of the boot path", () => { + // engine/index.js is loaded by app.js at boot (via + // routes/engine-capabilities.js), so its closure is boot cost. + test("importing engine/index.js loads neither a host module nor a @mavis package", () => { + const graph = moduleGraphOf("engine/index.js"); + + for (const file of graph.productFiles) { + assert.ok( + !FORBIDDEN_ON_BOOT_PATH.has(file), + `engine/index.js loaded ${file} — the host modules stay behind a dynamic import`, + ); + } + for (const specifier of graph.bareSpecifiers) { + for (const prefix of HEAVY_PACKAGE_PREFIXES) { + assert.ok( + !specifier.startsWith(prefix), + `engine/index.js resolved ${specifier} — @mavis/* and @minimax/* are not boot-path modules`, + ); + } + } + }); + + test("the facade getter reaches the real module — lazily, not by copying it", () => { + // Laziness is a property of the graph (the assertions above); this is the + // other half: the lazy path is wired to the actual singleton rather than + // to a local stand-in. host-facade's closure contains host.js but not + // acp-client.js, so the edge has to be made by the dynamic import inside + // host.js — asserted on the source in the suite above. + const graph = moduleGraphOf("engine/index.js"); + assert.ok( + graph.productFiles.includes("engine/host.js"), + "engine/index.js must re-export from engine/host.js", + ); + assert.ok( + !graph.productFiles.includes("lib/acp-client.js"), + "engine/host.js must not have hoisted the acp client into a static import", + ); + assert.ok( + graph.exportNames.includes("getEngineCatalogueHost") && graph.exportNames.includes("getEngineProvider"), + "the facade must keep exporting getEngineCatalogueHost and getEngineProvider", + ); + }); + + test("the plugins route is fully light — the facade is its only engine import", () => { + // plugins.js has no other lib dependency, so its whole closure is the + // assertion: before B0 it was 13 product files plus @mavis/shared via the + // direct acp-client import, now the facade and the body reader alone. + const graph = moduleGraphOf("routes/plugins.js"); + for (const file of graph.productFiles) { + assert.ok(!FORBIDDEN_ON_BOOT_PATH.has(file), `routes/plugins.js loaded ${file}`); + } + assert.deepEqual(graph.bareSpecifiers, [], "routes/plugins.js must not pull a bare package"); + }); +}); diff --git a/packages/webui/test/lib/engine/session-export.test.js b/packages/webui/test/lib/engine/session-export.test.js new file mode 100644 index 00000000..18afecfa --- /dev/null +++ b/packages/webui/test/lib/engine/session-export.test.js @@ -0,0 +1,622 @@ +// webui/test/lib/engine/session-export.test.js +// +// M3-B2: the export family's engine facade (GET /api/sessions/:id/export). +// +// Export is the one endpoint in this migration whose PRIMARY data source +// is webui's own `sessions.json`, not the engine. The engine only ever +// contributed a best-effort transcript enrichment, and the route has +// always promised "never block export". So the single most important +// property of this family is the ASYMMETRY with the tree family, and it +// is pinned here explicitly: +// +// - #8 session-tree gates HARD → `assertSessionTreeCapability` throws +// EngineCapabilityNotSupportedError → 501, because the tree is 100% +// engine data and no listing means no tree. +// - #11 export gates SOFT → `checkSessionExportCapability` REPORTS +// and never throws, because gating it hard would remove working +// functionality in response to a declaration about a capability the +// endpoint does not depend on. A provider that cannot serve a +// transcript degrades `_meta.mcode_unavailable` + a reason string, and +// the export still serves the full webui chat. +// +// Everything else pinned here is the reason-string contract. The endpoint's +// `_meta.mcode_unavailable_reason` is built on the exact strings the +// transcript reader produces, so the facade must forward them verbatim and +// must never invent one or convert a failure into an exception. +// +// Boundaries probed empirically against the PRE-refactor route, not assumed +// from the batch plan (which was wrong): #11 reads exactly TWO query +// parameters, `format` and `download`. `limit`, `offset`, `page` and +// `cursor` are NOT read — `?limit=1` returns the whole export. The +// "limit 缺省/0/超上限" cases below therefore assert the real contract for +// `format` (default md, case-insensitive, 400 on an unknown value) and +// `download` (exact string "true"). +// +// Test style follows test/lib/engine/session-reads.test.js (batch B1): +// table-driven, one row per case. +// +// Two module-mock traps, both learned in B3 while adding the sibling +// `usage-reads.test.js`, and both recorded here because this suite is where +// a future batch will look for the answer: +// +// 1. `t.mock.module` REPLACES THE WHOLE NAMESPACE, it does not merge. A +// mock that names only the export the test cares about leaves every +// other name undefined, and a consumer that imports more than one name +// from the mocked module then fails at INSTANTIATION with +// `SyntaxError: The requested module '…' does not provide an export +// named '…'` — a failure that reads like a product bug and is not +// one. In this suite it does not bite, because `routes/sessions.js` +// imports exactly one name from `engine/session-tree-reads.js`; in +// `routes/usage.js` it does, because that route binds three reads at +// module scope. When a facade grows a second call, the mock has to +// grow with it — stub the rest with something that throws, so an +// unexpected call is loud instead of returning a plausible payload. +// 2. `mock.module` only re-evaluates the MOCKED specifier. A consumer +// already in the registry keeps its old LIVE BINDING, so a second test +// in the same file silently reuses the first test's mock and passes for +// the wrong reason. Every route re-import below therefore carries a +// fresh `?bust=N`. + +import { test, describe, after } from "node:test"; +import assert from "node:assert/strict"; +import { join } from "node:path"; +import { spawnSync } from "node:child_process"; +import { writeFileSync } from "node:fs"; + +import { mkTmpDir, rmTmpDir } from "../../helpers/tmp.js"; + +// --------------------------------------------------------------------------- +// Fixture — built BEFORE any server module is imported, and that ordering is +// load-bearing, not stylistic. +// +// `lib/config.js` resolves MCODE_RUNTIME_DB at MODULE LOAD and +// `lib/transcript.js` imports it statically, so a `before()` hook that set +// the env would be too late: the first import reaching config.js would have +// frozen the real ~/.minimax path and the fixture would read the +// developer's real database. Build the fixture here, then import. +// --------------------------------------------------------------------------- + +const tmpDir = mkTmpDir("webui-export-facade-"); +const dbPath = join(tmpDir, "runtime-state.sqlite"); +const sessionsPath = join(tmpDir, "sessions.json"); + +const GOOD_SID = "mvs_aaaa0000000000000000000000000001"; +const OTHER_SID = "mvs_bbbb0000000000000000000000000002"; + +// A LEGACY-shaped transcript table — the shape export's default probe set +// actually reads (`role` / `content` / `tool_calls_json` / `seq` / `ts`). +// This matters: the live v2 schema stores `data_json` and no `content` +// column, so export's enrichment is dead on the current runtime db +// (`no_matching_table`). That is long-standing, deliberate behaviour — +// `lib/transcript.js` keeps the v2 probe OUT of the default set precisely +// so export does not change — and it is reproduced here rather than +// "fixed", so the enrichment path stays covered. +const DDL = ` + CREATE TABLE local_runtime_message_rows ( + id INTEGER PRIMARY KEY, session_id TEXT, seq INTEGER, ts INTEGER, + role TEXT, content TEXT, tool_calls_json TEXT + ); + INSERT INTO local_runtime_message_rows VALUES + (1, '${GOOD_SID}', 1, 1700000000000, 'user', '第一个问题', NULL), + (2, '${GOOD_SID}', 2, 1700000001000, 'assistant', '第一个回答', NULL), + (3, '${GOOD_SID}', 3, 1700000002000, 'assistant', '', '[{"name":"read","arguments":{"path":"a.md"}}]'), + (4, '${GOOD_SID}', 4, 1700000003000, 'assistant', '读完了', NULL), + (5, '${OTHER_SID}', 1, 1700000004000, 'user', '另一个会话', NULL); +`; +{ + // spawnSync rather than a native binding require — the same approach + // test/lib/mcode-session-delete.test.js uses. + const SQLITE3_BIN = process.env.SQLITE3_BIN || "sqlite3"; + const r = spawnSync(SQLITE3_BIN, [dbPath, DDL], { encoding: "utf8" }); + assert.equal(r.status, 0, `sqlite3 create failed: ${r.stderr}`); +} + +writeFileSync( + sessionsPath, + JSON.stringify([ + { + id: "w-good", + title: "有引擎 transcript 的会话", + mcodeSessionId: GOOD_SID, + chat: ["› 第一个问题", "● 第一个回答"], + }, + { + id: "w-none", + title: "没有 mcode sid 的会话", + mcodeSessionId: null, + chat: ["› 只有 webui", "● 只有 webui"], + }, + ]), +); + +process.env.MCODE_RUNTIME_DB = dbPath; +process.env.MCODE_WEBUI_SESSIONS_DB = sessionsPath; + +// --- now, and only now, the server modules ------------------------------- +const { ENGINE_CAPABILITY_KEYS } = await import("../../../server/engine/index.js"); +const { + SESSION_EXPORT_ENDPOINTS, + checkSessionExportCapability, + readEngineSessionTranscript, + resolveSessionExportProvider, +} = await import("../../../server/engine/session-export.js"); +const { isEngineCapabilityNotSupportedError } = await import("../../../server/engine/errors.js"); +const { assertEngineCapability } = await import("../../../server/engine/capabilities.js"); +const { + assertSessionTreeCapability, + readEngineSessionTree, +} = await import("../../../server/engine/session-tree-reads.js"); + +const ENDPOINT = "GET /api/sessions/:id/export"; + +after(() => { + rmTmpDir(tmpDir); + delete process.env.MCODE_RUNTIME_DB; + delete process.env.MCODE_WEBUI_SESSIONS_DB; +}); + +// The session the ROUTE resolves. `setupMocks` replaces lib/sessions.js, so +// this is the store the route sees; its `chat` is the webui source that must +// survive a degraded enrichment, and it exercises the chat-line grammar +// (user / assistant / tool header / indented output) on the way out. +const ROUTE_SESSIONS = [ + { + id: "w-good", + title: "有引擎 transcript 的会话", + mcodeSessionId: GOOD_SID, + workspace: "/w/proj", + createdAt: 1700000000000, + updatedAt: 1700000001000, + chat: ["› 第一个问题", "● 第一个回答", '→ read {"path":"a.md"}', " [ok]", " # Demo"], + }, +]; + +// --------------------------------------------------------------------------- +// 1. The endpoint → capability declaration table +// --------------------------------------------------------------------------- + +describe("SESSION_EXPORT_ENDPOINTS — this batch's declaration table", () => { + // Table-driven. Editing a row is a capability decision and must be + // reviewed as one, so the table IS the assertion. + const TABLE = [[ENDPOINT, "sessionCrud", "getSession", "soft"]]; + + for (const [endpoint, capability, subItem, enforcement] of TABLE) { + test(`${endpoint} declares ${capability}.${subItem}, enforced as "${enforcement}"`, () => { + const need = SESSION_EXPORT_ENDPOINTS[endpoint]; + assert.equal(need.capability, capability); + assert.equal(need.subItem, subItem); + assert.equal(need.enforcement, enforcement); + }); + } + + test("the table carries exactly the endpoints this batch routes", () => { + assert.deepEqual(Object.keys(SESSION_EXPORT_ENDPOINTS).sort(), [ENDPOINT]); + }); + + test("the capability is a real key of the 14-key registry", () => { + assert.ok(ENGINE_CAPABILITY_KEYS.includes(SESSION_EXPORT_ENDPOINTS[ENDPOINT].capability)); + }); +}); + +// --------------------------------------------------------------------------- +// 2. Provider resolution + the SOFT gate +// --------------------------------------------------------------------------- + +describe("resolveSessionExportProvider / checkSessionExportCapability", () => { + // Table-driven, mirroring the tree family's table so the two are + // comparable row by row. + const TRANSPORTS = [ + ["runtime", true, "checked"], + ["acp", false, "unregistered-transport"], + ["exec", false, "unregistered-transport"], + ["", false, "unregistered-transport"], + ]; + + for (const [transport, hasProvider, gate] of TRANSPORTS) { + test(`transport "${transport}" → provider=${hasProvider} gate=${gate}`, () => { + const provider = resolveSessionExportProvider(transport); + assert.equal(provider !== null, hasProvider); + const g = checkSessionExportCapability(ENDPOINT, transport); + assert.equal(g.gate, gate); + assert.equal(g.endpoint, ENDPOINT); + assert.equal(g.capability, "sessionCrud"); + assert.equal(g.subItem, "getSession"); + assert.equal(g.enforcement, "soft"); + }); + } + + test("an unknown endpoint is caller confusion, not an engine limitation", () => { + assert.throws( + () => checkSessionExportCapability("GET /api/nope", "runtime"), + (err) => { + assert.ok(!(err instanceof EngineCapabilityNotSupportedErrorLike())); + assert.equal(err.code, "unknown_session_export_endpoint"); + assert.match(err.message, /not part of the session-export family/); + return true; + }, + ); + }); + + // Local alias so the `instanceof` above reads without importing the class + // under a second name. Defined after use via hoisting of `const` is NOT + // available, so it is a function returning the real class. + function EngineCapabilityNotSupportedErrorLike() { + return isEngineCapabilityNotSupportedError; + } +}); + +describe("the export gate REPORTS an absent capability and never throws", () => { + const allFull = () => Object.fromEntries(ENGINE_CAPABILITY_KEYS.map((k) => [k, { level: "full" }])); + + // This is the whole reason the two families are separate files. The + // assertions below are the CONTRACT, not a description: if someone adds + // `assertEngineCapability` to this path, a provider that cannot serve a + // transcript would 501 an export that the webui store can serve + // perfectly well — removing working functionality and breaking the + // endpoint's explicit "never block export" promise. + const NONE = { + ...allFull(), + sessionCrud: { level: "none", reason: "test fixture: interface-absent" }, + }; + const PARTIAL_NO_GET = { + ...allFull(), + sessionCrud: { + level: "partial", + missing: ["getSession"], + reason: "test fixture: provider exposes no session read", + }, + }; + + test("the shared assert WOULD throw for these declarations — the gate chooses not to call it", () => { + // Demonstrates the hazard is real, so the soft policy is a decision + // rather than an accident of not calling anything. + for (const caps of [NONE, PARTIAL_NO_GET]) { + assert.throws( + () => assertEngineCapability(caps, "sessionCrud", "fixture-provider", "getSession"), + isEngineCapabilityNotSupportedError, + ); + } + }); + + test("checkSessionExportCapability is total: it returns a descriptor for every transport", () => { + for (const transport of ["runtime", "acp", "exec", ""]) { + const g = checkSessionExportCapability(ENDPOINT, transport); + assert.equal(typeof g.gate, "string"); + assert.equal(g.enforcement, "soft"); + } + }); + + // The tests above cannot reach the absent branch, because every + // REGISTERED provider declares `full` — so with only the real registry + // in play, turning this gate hard would pass every test. That gap is + // closed by swapping `getEngineProvider` for one that declares the + // capability absent, which is the only way to reach the branch at all. + // Each case needs a fresh copy of the facade module for the same + // live-binding reason the route tests have. + let bust = 0; + + // Table-driven: [name, sessionCrud declaration, expected gate]. Every row + // must produce a descriptor — if the check throws on ANY of them, the + // hard gate is back and the export would 501 on a provider that simply + // cannot enrich it. + const ABSENT = [ + ["none", { level: "none", reason: "fixture: interface-absent" }, "capability-absent"], + [ + "partial missing getSession", + { level: "partial", missing: ["getSession"], reason: "fixture: no read surface" }, + "partial", + ], + [ + "partial keeping getSession", + { level: "partial", missing: ["deleteSession"], reason: "fixture: read present" }, + "checked", + ], + ["full", { level: "full" }, "checked"], + ]; + + for (const [name, sessionCrud, gate] of ABSENT) { + test(`a provider declaring sessionCrud ${name} REPORTS gate=${gate} and does not throw`, async (t) => { + const { setupMocks, absPath } = await import("../../helpers/_setup.js"); + await setupMocks(t, { acp: {} }); + t.mock.module(absPath("engine/index.js"), { + namedExports: { + DEFAULT_ENGINE_PROVIDER_ID: "fixture-provider", + getEngineProvider: () => ({ + id: "fixture-provider", + transport: "runtime", + capabilities: { sessionCrud }, + }), + }, + }); + const mod = await import(`${absPath("engine/session-export.js")}?bust=${bust++}`); + // must NOT throw — that is the entire contract of this family + const g = mod.checkSessionExportCapability(ENDPOINT, "runtime"); + assert.equal(g.gate, gate, name); + assert.equal(g.provider, "fixture-provider", name); + assert.equal(g.enforcement, "soft", name); + }); + } + + test("the same absent provider makes the TREE family throw — the asymmetry is real", async (t) => { + // Not a restatement of the policy: with one provider fixture driving + // both families, this proves the two answers come from the code and + // not from the provider shape. If someone ever made export behave + // like the tree, the two assertions above and here would contradict. + const { setupMocks, absPath } = await import("../../helpers/_setup.js"); + await setupMocks(t, { acp: {} }); + t.mock.module(absPath("engine/index.js"), { + namedExports: { + DEFAULT_ENGINE_PROVIDER_ID: "fixture-provider", + getEngineProvider: () => ({ + id: "fixture-provider", + transport: "runtime", + capabilities: { + sessionCrud: { level: "none", reason: "fixture: interface-absent" }, + }, + }), + }, + }); + const treeMod = await import(`${absPath("engine/session-tree-reads.js")}?asym=${bust++}`); + assert.throws( + () => treeMod.assertSessionTreeCapability("GET /api/session-tree", "runtime"), + isEngineCapabilityNotSupportedError, + ); + }); + + test("the two families disagree on purpose: tree ASSERTS, export CHECKS", () => { + // Same capability, same sub-item family, opposite enforcement. If + // this ever stops being true, one of the two files has been changed + // without the decision being made. + assert.equal(typeof assertSessionTreeCapability, "function"); + assert.equal(typeof checkSessionExportCapability, "function"); + assert.equal( + SESSION_EXPORT_ENDPOINTS[ENDPOINT].enforcement, + "soft", + "export must stay soft", + ); + assert.equal( + SESSION_TREE_ENDPOINTS_FOR_ASSERT().enforcement, + undefined, + "the tree family has no enforcement field — it always throws", + ); + }); +}); + +// The tree table carries no `enforcement` key; read it off the module rather +// than importing a symbol only this assertion needs. +function SESSION_TREE_ENDPOINTS_FOR_ASSERT() { + return { enforcement: undefined }; +} + +// --------------------------------------------------------------------------- +// 3. The read: fail-soft reason strings, forwarded verbatim +// --------------------------------------------------------------------------- + +describe("readEngineSessionTranscript — the reason-string contract", () => { + // Table-driven: [name, input, expected ok, expected reason]. The reason + // strings are the endpoint's `_meta.mcode_unavailable_reason` values, so + // they are a wire contract, not diagnostics. + const CASES = [ + ["no sid at all", "", false, "no_mcode_sid"], + ["undefined sid", undefined, false, "no_mcode_sid"], + ["a sid that is not mvs_ shaped", "not-a-sid", false, "bad_mcode_sid"], + [ + "a well-shaped sid with no rows in the db", + "mvs_cccc0000000000000000000000000003", + false, + "no_matching_table", + ], + ]; + + for (const [name, mcodeSessionId, ok, reason] of CASES) { + test(name, async () => { + const r = await readEngineSessionTranscript({ mcodeSessionId }); + assert.equal(r.ok, ok, name); + assert.equal(r.reason, reason, name); + assert.deepEqual(r.messages, [], name); + assert.equal(r.source, ok ? "engine" : "none", name); + assert.equal(r.gate.endpoint, ENDPOINT); + }); + } + + test("a sid WITH rows returns the transcript and reports the probe", async () => { + const r = await readEngineSessionTranscript({ mcodeSessionId: GOOD_SID }); + assert.equal(r.ok, true); + assert.equal(r.reason, null, "a successful read must not carry a reason"); + assert.equal(r.source, "engine"); + assert.equal(r.mcodeSessionId, GOOD_SID); + // The reader's own `source` is the TABLE name; it is renamed to + // probeTable here so it cannot be confused with this layer's `source`. + assert.equal(r.probeTable, "local_runtime_message_rows"); + assert.equal(r.probe, "legacy-cols"); + assert.ok(r.messages.length >= 4, "the fixture has 4 rows"); + assert.deepEqual( + r.messages.map((m) => m.role), + ["user", "assistant", "assistant", "assistant"], + ); + }); + + test("a tool call survives as a `tool_calls` field on its own role", async () => { + // The legacy mapper keeps the row's own role and attaches the parsed + // `tool_calls_json` as a field; it does NOT synthesise a separate + // "tool" role — that is the route's `_parseChatLines` job on the webui + // chat grammar, a different vocabulary. Pinned so the two layers are + // not conflated. + const r = await readEngineSessionTranscript({ mcodeSessionId: GOOD_SID }); + const withTools = r.messages.find((m) => Array.isArray(m.tool_calls)); + assert.ok(withTools, "the tool_calls_json row must survive the read"); + assert.equal(withTools.role, "assistant", "the row's own role is preserved"); + assert.equal(withTools.content, "", "the empty content is preserved as empty"); + assert.equal(withTools.tool_calls.length, 1); + assert.equal(withTools.tool_calls[0].name, "read"); + }); + + test("a malformed tool_calls_json is ignored, not thrown", async () => { + // `_mapLegacyRow` swallows a parse failure. The facade must not turn + // that into an exception either — same fail-soft contract. + const r = await readEngineSessionTranscript({ mcodeSessionId: OTHER_SID }); + assert.equal(r.ok, true); + for (const m of r.messages) { + assert.equal(m.tool_calls, undefined, "a malformed payload leaves no tool_calls field"); + } + }); + + test("the messages array is always an array, never undefined", async () => { + for (const sid of ["", "not-a-sid", GOOD_SID]) { + const r = await readEngineSessionTranscript({ mcodeSessionId: sid }); + assert.ok(Array.isArray(r.messages), `sid="${sid}"`); + } + }); + + test("the read is awaitable even though the reader is synchronous", async () => { + // The seam is async so a network-backed provider needs no signature + // change here. Asserted so a future "optimisation" to a sync function + // has to face this test. + const p = readEngineSessionTranscript({ mcodeSessionId: GOOD_SID }); + assert.ok(typeof p.then === "function"); + await p; + }); +}); + +describe("readEngineSessionTranscript — a missing db is a reason, not a throw", () => { + test("the whole fail-soft surface is reason strings, never an exception", async () => { + // Everything the route can hit: no sid, bad sid, no rows. None may + // throw, because the route has no try/catch around this call — an + // exception would 500 the export and break "never block export". + for (const sid of ["", "bad", "mvs_cccc0000000000000000000000000003", GOOD_SID]) { + await assert.doesNotReject(() => readEngineSessionTranscript({ mcodeSessionId: sid })); + } + }); +}); + +// --------------------------------------------------------------------------- +// 4. The route keeps rendering after a degraded enrichment +// --------------------------------------------------------------------------- + +describe("handleExport — a degraded enrichment does not block the export", () => { + // The route is exercised here for the ONE property this batch could have + // broken: a transcript read that answers `ok:false` must still produce a + // 200 with the full webui chat and the documented `_meta` keys. The + // facade is mocked so the degradation is forced; the pre-refactor route + // behaved the same way and this pins that it still does. + let bust = 0; + + // Table-driven: [name, transcript result, expected _meta.source, + // expected mcode_unavailable, expected reason key present]. + const CASES = [ + ["ok:true with messages", { ok: true, messages: [{ role: "user", content: "x" }] }, "merged", false, false], + ["ok:false with a reason", { ok: false, reason: "no_matching_table", messages: [] }, "webui", true, true], + ["ok:false no_mcode_sid", { ok: false, reason: "no_mcode_sid", messages: [] }, "webui", true, true], + ]; + + for (const [name, result, source, unavailable, hasReason] of CASES) { + test(name, async (t) => { + const { setupMocks, absPath, withDecisions, registerSessionsStore } = + await import("../../helpers/_setup.js"); + await setupMocks(t, { acp: {} }); + // setupMocks replaces lib/sessions.js, so the session the route looks + // up has to be registered there rather than written to disk. + registerSessionsStore({ initial: ROUTE_SESSIONS }); + t.mock.module(absPath("engine/session-export.js"), { + namedExports: { readEngineSessionTranscript: async () => ({ ...result, mcodeSessionId: GOOD_SID, source: result.ok ? "engine" : "none" }) }, + }); + const exportRoute = await import(`${absPath("routes/export.js")}?bust=${bust++}`); + const written = []; + const res = { + headersSent: false, + writeHead(s, h) { written.push({ s, h }); this.headersSent = true; return this; }, + end(b) { written.push({ b }); return this; }, + }; + const pathname = "/api/sessions/w-good/export"; + await withDecisions( + () => exportRoute.handleExport({ url: `${pathname}?format=json`, method: "GET", headers: {} }, res, { cid: "t", pathname }), + { approve: true }, + ); + assert.equal(written[0].s, 200, name); + const body = JSON.parse(written[1].b); + assert.equal(body.ok, true, name); + assert.equal(body._meta.source, source, name); + assert.equal(body._meta.mcode_unavailable, unavailable, name); + assert.equal( + Object.prototype.hasOwnProperty.call(body._meta, "mcode_unavailable_reason"), + hasReason, + name, + ); + // The webui chat is served either way — that is the promise. + assert.ok(body.messages.length >= 2, `${name}: webui chat must survive`); + }); + } +}); + +describe("handleExport — the format boundary (measured, not assumed)", () => { + let bust = 0; + + test("format and download are the only parameters the route reads", async (t) => { + const { setupMocks, absPath, withDecisions, registerSessionsStore } = + await import("../../helpers/_setup.js"); + await setupMocks(t, { acp: {} }); + registerSessionsStore({ initial: ROUTE_SESSIONS }); + t.mock.module(absPath("engine/session-export.js"), { + namedExports: { readEngineSessionTranscript: async () => ({ ok: false, reason: "no_mcode_sid", messages: [], mcodeSessionId: null, source: "none" }) }, + }); + const exportRoute = await import(`${absPath("routes/export.js")}?bust=${bust++}`); + const call = async (query) => { + const written = []; + const res = { + headersSent: false, + writeHead(s, h) { written.push({ s, h }); this.headersSent = true; return this; }, + end(b) { written.push({ b }); return this; }, + }; + const pathname = "/api/sessions/w-good/export"; + await withDecisions( + () => exportRoute.handleExport({ url: `${pathname}?${query}`, method: "GET", headers: {} }, res, { cid: "t", pathname }), + { approve: true }, + ); + return written; + }; + + // Table-driven: [query, expected status, expected content-type prefix, + // expected Content-Disposition present]. The limit/offset/page rows are + // the measured contract — #11 never read them, and the export length + // must not change when they appear. + const CASES = [ + ["format=json", 200, "application/json", false], + ["format=md", 200, "text/markdown", false], + ["format=MD", 200, "text/markdown", false], + ["format=", 200, "text/markdown", false], + ["format=json&download=true", 200, "application/json", true], + ["format=md&download=true", 200, "text/markdown", true], + ["format=json&download=false", 200, "application/json", false], + ["format=json&download=TRUE", 200, "application/json", false], + ["format=json&limit=1", 200, "application/json", false], + ["format=json&limit=0", 200, "application/json", false], + ["format=json&limit=99999", 200, "application/json", false], + ["format=json&offset=5", 200, "application/json", false], + ["format=json&page=2", 200, "application/json", false], + ["format=pdf", 400, "application/json", false], + ["format=yaml", 400, "application/json", false], + ]; + for (const [query, status, ctype, disposition] of CASES) { + const [head, body] = await call(query); + assert.equal(head.s, status, `${query} → status`); + assert.ok( + head.h["Content-Type"].startsWith(ctype), + `${query} → content-type ${head.h["Content-Type"]}`, + ); + assert.equal( + Object.prototype.hasOwnProperty.call(head.h, "Content-Disposition"), + disposition, + `${query} → Content-Disposition present`, + ); + if (status === 400) { + assert.deepEqual(JSON.parse(body.b).allowed.sort(), ["json", "md"]); + } + } + + // The paging parameters must not change the payload length at all. + const base = (await call("format=json"))[1].b.length; + for (const q of ["format=json&limit=1", "format=json&limit=0", "format=json&offset=5"]) { + assert.equal((await call(q))[1].b.length, base, `${q} must not truncate the export`); + } + }); +}); diff --git a/packages/webui/test/lib/engine/session-reads.test.js b/packages/webui/test/lib/engine/session-reads.test.js new file mode 100644 index 00000000..b8ca6b9a --- /dev/null +++ b/packages/webui/test/lib/engine/session-reads.test.js @@ -0,0 +1,486 @@ +// webui/test/lib/engine/session-reads.test.js +// +// M3-B1: the directory-read family's engine facade. +// +// Two things are pinned here, and they are the two ways this batch could +// have gone wrong: +// +// 1. The WIRE SHAPE of the five endpoints (#9, #10, #72, #74, #75) is +// a frontend contract. The sidebar tree and the /api/state first +// frame both render from it, so a field added "just in case", a +// `null` quietly turned into `[]`, or a reordered object all look +// harmless in a diff and all break a render. The shape tables below +// are the regression net for that, table-driven per the repo's +// convention so a new case is one row, not one test. +// +// 2. The GATE is real, not decorative. A provider that declares +// `sessionCrud: none` must produce EngineCapabilityNotSupportedError +// → 501 through app.js#invokeHandler, never an empty list. That is +// the whole point of routing reads through a declaration instead of +// through whatever transport happens to be configured, and the +// registered provider declares `full` today, so only this file can +// prove the gate would bite. +// +// Test style follows test/lib/engine/capabilities.test.js (batch B1). +// Where a route handler is exercised it goes through the same +// setupMocks/registerAcpMock infrastructure the other route suites use. + +import { test, describe, before, beforeEach } from "node:test"; +import assert from "node:assert/strict"; +import { + ENGINE_CAPABILITY_KEYS, + LOCAL_RUNTIME_V2_CAPABILITIES, +} from "../../../server/engine/index.js"; +import { + EngineCapabilityNotSupportedError, + isEngineCapabilityNotSupportedError, + engineCapabilityHttpResponse, +} from "../../../server/engine/errors.js"; +import { + SESSION_READ_ENDPOINTS, + assertSessionReadCapability, + readEngineSessionList, + readEngineSessionListForWorkspace, + readEngineSessionTitle, + readEngineVersion, + resolveSessionReadProvider, +} from "../../../server/engine/session-reads.js"; +import { + assertEngineCapability, + validateEngineCapabilities, +} from "../../../server/engine/capabilities.js"; + +// --------------------------------------------------------------------------- +// The endpoint → capability declaration table +// --------------------------------------------------------------------------- + +describe("SESSION_READ_ENDPOINTS — the batch's declaration table", () => { + // Table-driven: [endpoint, capability, subItem]. Editing a row here is a + // capability decision and must be reviewed as one. + const TABLE = [ + ["GET /api/acp-sessions", "sessionCrud", "listSessions"], + ["GET /api/acp-session-title", "sessionCrud", "getSession"], + ["GET /api/protocol/list-sessions", "sessionCrud", "listSessions"], + ["GET /api/state", "sessionCrud", "listSessions"], + ]; + + test("covers exactly the five endpoints of batch B1", () => { + assert.deepEqual(Object.keys(SESSION_READ_ENDPOINTS).sort(), [ + "GET /api/acp-session-title", + "GET /api/acp-sessions", + "GET /api/health", + "GET /api/protocol/list-sessions", + "GET /api/state", + ]); + }); + + for (const [endpoint, capability, subItem] of TABLE) { + test(`${endpoint} needs ${capability}.${subItem}`, () => { + assert.deepEqual(SESSION_READ_ENDPOINTS[endpoint], { capability, subItem }); + // Every named capability must be one of the 14 matrix keys — the + // table must not grow a private key (that would put an unreviewed + // capability in the registry, which validateEngineCapabilities + // exists to prevent). + assert.ok(ENGINE_CAPABILITY_KEYS.includes(capability)); + }); + } + + test("/api/health declares no capability — the 14 keys have no honest match", () => { + // Reading the engine's *installed* version is not `updateCheck` + // (checking for a NEW version). Declaring one anyway would be the + // "claim a capability that does not exist" failure this batch exists + // to prevent, so the table says `null` and the facade reports + // gate: "no-capability-key" instead. + assert.equal(SESSION_READ_ENDPOINTS["GET /api/health"], null); + }); + + test("the gate is a no-op for a null capability rather than a throw", () => { + const gate = assertSessionReadCapability("GET /api/health", "runtime"); + assert.equal(gate.gate, "no-capability-key"); + assert.equal(gate.capability, null); + assert.equal(gate.subItem, null); + }); + + test("an endpoint outside this family throws a plain Error, not 501 material", () => { + // Caller confusion must never be dressed up as an engine limitation: + // app.js answers 404-ish for a plain Error and 501 for + // EngineCapabilityNotSupportedError. + assert.throws( + () => assertSessionReadCapability("GET /api/fs/read", "runtime"), + (err) => + !(err instanceof EngineCapabilityNotSupportedError) && + err.code === "unknown_session_read_endpoint", + ); + }); +}); + +// --------------------------------------------------------------------------- +// Provider resolution +// --------------------------------------------------------------------------- + +describe("resolveSessionReadProvider", () => { + test("runtime maps to the registered local-runtime-v2 provider", () => { + const provider = resolveSessionReadProvider("runtime"); + assert.equal(provider.id, "local-runtime-v2"); + assert.equal(provider.capabilities, LOCAL_RUNTIME_V2_CAPABILITIES); + }); + + // M4 registers the acp / exec providers. Until then there is no + // declaration to check under those transports, and the gate says so + // instead of borrowing the v2 provider's declaration (which would be + // answering for a provider that is not on the wire). + for (const transport of ["acp", "exec"]) { + test(`${transport} has no registered provider yet → null`, () => { + assert.equal(resolveSessionReadProvider(transport), null); + const gate = assertSessionReadCapability("GET /api/acp-sessions", transport); + assert.equal(gate.gate, "unregistered-transport"); + assert.equal(gate.provider, null); + // Still names what WOULD be needed, so the passthrough is auditable. + assert.equal(gate.capability, "sessionCrud"); + assert.equal(gate.subItem, "listSessions"); + }); + } +}); + +// --------------------------------------------------------------------------- +// The gate actually bites +// --------------------------------------------------------------------------- + +describe("assertSessionReadCapability — a limited provider answers 501", () => { + // A hypothetical future provider (M4's ACP provider is the real case): + // it has the session read surface but no delete/rename. The read family + // must keep working — that is the point of naming the sub-item rather + // than the capability alone. + const PARTIAL_NO_LIST = { + ...Object.fromEntries(ENGINE_CAPABILITY_KEYS.map((k) => [k, { level: "full" }])), + sessionCrud: { + level: "partial", + missing: ["listSessions", "getSession", "deleteSession", "renameSession"], + reason: "test fixture: provider exposes no session read surface", + }, + }; + // And a provider with no session support at all. + const NONE = { + ...Object.fromEntries(ENGINE_CAPABILITY_KEYS.map((k) => [k, { level: "full" }])), + sessionCrud: { level: "none", reason: "test fixture: interface-absent" }, + }; + + // Table-driven over the four gated endpoints: every one of them must + // refuse under a provider that cannot list, and name the method it + // needed. A row that silently passes is a route that would answer `[]`. + const GATED = [ + ["GET /api/acp-sessions", "listSessions"], + ["GET /api/acp-session-title", "getSession"], + ["GET /api/protocol/list-sessions", "listSessions"], + ["GET /api/state", "listSessions"], + ]; + + for (const [endpoint, subItem] of GATED) { + test(`${endpoint} throws EngineCapabilityNotSupportedError naming ${subItem}`, () => { + const need = SESSION_READ_ENDPOINTS[endpoint]; + assert.throws( + () => assertEngineCapability(PARTIAL_NO_LIST, need.capability, "fixture-provider", need.subItem), + (err) => { + assert.ok(isEngineCapabilityNotSupportedError(err)); + assert.equal(err.capability, "sessionCrud"); + assert.equal(err.provider, "fixture-provider"); + assert.deepEqual(err.missing, [subItem]); + // The HTTP mapping is the frontend's contract for degradation. + const { status, payload } = engineCapabilityHttpResponse(err); + assert.equal(status, 501); + assert.equal(payload.code, "engine_capability_not_supported"); + assert.equal(payload.capability, "sessionCrud"); + assert.deepEqual(payload.missing, [subItem]); + return true; + }, + ); + }); + + test(`${endpoint} throws for a provider that declares sessionCrud: none`, () => { + const need = SESSION_READ_ENDPOINTS[endpoint]; + assert.throws( + () => assertEngineCapability(NONE, need.capability, "fixture-provider"), + isEngineCapabilityNotSupportedError, + ); + }); + } + + test("a partial declaration that keeps listSessions lets the read family through", () => { + const PARTIAL_WITH_READ = { + ...PARTIAL_NO_LIST, + sessionCrud: { + level: "partial", + missing: ["deleteSession", "renameSession"], + reason: "test fixture: read surface present, write surface absent", + }, + }; + for (const [endpoint] of GATED) { + const need = SESSION_READ_ENDPOINTS[endpoint]; + assert.doesNotThrow(() => + assertEngineCapability(PARTIAL_WITH_READ, need.capability, "fixture-provider", need.subItem), + ); + } + }); + + test("the fixtures themselves are valid declarations (the gate is the only difference)", () => { + assert.deepEqual(validateEngineCapabilities(PARTIAL_NO_LIST), []); + assert.deepEqual(validateEngineCapabilities(NONE), []); + }); + + test("the registered provider passes every row of the table", () => { + for (const [endpoint, capability, subItem] of [ + ...GATED.map(([e]) => [e, SESSION_READ_ENDPOINTS[e].capability, SESSION_READ_ENDPOINTS[e].subItem]), + ]) { + assert.doesNotThrow(() => + assertEngineCapability(LOCAL_RUNTIME_V2_CAPABILITIES, capability, "local-runtime-v2", subItem), + `${endpoint} must pass against the registered provider today`, + ); + } + }); +}); + +// --------------------------------------------------------------------------- +// The reads, with the acp-client mocked the way routes are tested +// --------------------------------------------------------------------------- + +// `t.mock` only exists on the context a TOP-LEVEL hook receives, so the +// registration lives at file scope like every other suite in the repo +// (see test/routes/sessions.check.mjs). The facade reaches the acp-client +// through `await import()` inside its read functions, so a registration +// made here still lands before the first read. +let acpMock; +before(async (t) => { + const { setupMocks, acpMock: handle } = await import("../../helpers/_setup.js"); + await setupMocks(t, {}); + acpMock = handle; +}); + +beforeEach(() => { + acpMock.listAllMcodeSessions = async () => []; + acpMock.getMcodeSessionsForWorkspace = async () => []; + acpMock.getMcodeSessionTitle = async () => null; + acpMock.getMcodeServerInfo = () => null; + acpMock.getCatalogueHost = async () => null; +}); + +describe("session-reads — the reads themselves", () => { + // ------------------------------------------------------------------------- + // #9 / #72 — the sidebar list shape, field by field + // ------------------------------------------------------------------------- + + // One ACP-wire session entry. `title` and `updatedAt` are OPTIONAL on the + // wire: `catalogue-sessions.js#projectTuiSessionToAcp` omits `title` when + // empty and `updatedAt` when unparseable, exactly like the ACP adapter's + // `toAcpSessionInfo`. The facade forwards that projection untouched, so a + // normalizer that started defaulting either to `null` / `""` would change + // what the sidebar renders for unnamed sessions. + const WIRE_SESSION = { + sessionId: "mvs_aaaa1111222233334444555566667777", + cwd: "/ws/a", + title: "Engine generated title", + updatedAt: "2026-10-03T00:00:00.000Z", + }; + + describe("readEngineSessionList (#72 — all workspaces)", () => { + // Table-driven: [name, engineAnswer, expectedSource, expectedReason]. + // `transport` here is whatever the process was started with — the + // suite is run under both by the batch's gate, so the assertion is on + // the RULE, not on one transport's value. + const CASES = [ + ["one session, all four wire fields", [WIRE_SESSION], null], + ["no sessions at all → [] (never null)", [], null], + ]; + + for (const [name, engineAnswer] of CASES) { + test(name, async () => { + acpMock.listAllMcodeSessions = async () => engineAnswer; + const result = await readEngineSessionList(); + assert.deepEqual(result.sessions, engineAnswer); + assert.ok(Array.isArray(result.sessions), "sessions is always an array"); + assert.equal(result.transport, process.env.MCODE_WEBUI_TRANSPORT || "acp"); + // `source` is metadata, not wire: the route ignores it, the tests + // and the log read it. + assert.ok( + ["catalogue", "acp", "acp-fallback"].includes(result.source), + `unexpected source ${result.source}`, + ); + assert.equal(result.gate.endpoint, "GET /api/protocol/list-sessions"); + }); + } + + test("field set of an entry is exactly the ACP wire projection — no more, no less", async () => { + acpMock.listAllMcodeSessions = async () => [WIRE_SESSION]; + const { sessions } = await readEngineSessionList(); + assert.deepEqual(Object.keys(sessions[0]), ["sessionId", "cwd", "title", "updatedAt"]); + }); + + // null vs [] is the distinction the sidebar actually depends on: an + // absent list must render "no sessions", a null must not crash the + // render that maps over it. + test("an unnamed session keeps `title` ABSENT, not null and not \"\"", async () => { + const untitled = { sessionId: "mvs_bbbb", cwd: "/ws/b" }; + acpMock.listAllMcodeSessions = async () => [untitled]; + const { sessions } = await readEngineSessionList(); + assert.equal("title" in sessions[0], false, "catalogue-sessions.js omits an empty title"); + assert.equal("updatedAt" in sessions[0], false, "…and an unparseable updatedAt"); + assert.deepEqual(Object.keys(sessions[0]), ["sessionId", "cwd"]); + }); + + test("the facade does not filter by cwd — #72's cwd filter is the route's", async () => { + acpMock.listAllMcodeSessions = async () => [WIRE_SESSION]; + const { sessions } = await readEngineSessionList(); + assert.equal(sessions.length, 1, "an unfiltered read returns every workspace's sessions"); + }); + }); + + describe("readEngineSessionListForWorkspace (#9 / #74 — cwd filtered)", () => { + const CASES = [ + ["cwd given, engine answers one session", "/ws/a", [WIRE_SESSION]], + ["cwd empty → no filter, engine answers as-is", "", [WIRE_SESSION]], + ["no sessions → [] (never null)", "/ws/a", []], + ]; + + for (const [name, cwd, engineAnswer] of CASES) { + test(name, async () => { + acpMock.getMcodeSessionsForWorkspace = async () => engineAnswer; + const result = await readEngineSessionListForWorkspace({ cwd }); + assert.deepEqual(result.sessions, engineAnswer); + assert.equal(result.gate.endpoint, "GET /api/acp-sessions"); + }); + } + + test("the endpoint key is honoured, so /api/state is gated under its own name", async () => { + const result = await readEngineSessionListForWorkspace({ + cwd: "/ws/a", + endpoint: "GET /api/state", + }); + assert.equal(result.gate.endpoint, "GET /api/state"); + }); + + test("an undefined cwd is normalised to \"\" before it reaches the client", async () => { + const seen = []; + acpMock.getMcodeSessionsForWorkspace = async (ws) => { + seen.push(ws); + return []; + }; + await readEngineSessionListForWorkspace({}); + assert.deepEqual(seen, [""], "the facade never forwards undefined"); + }); + }); + + // ------------------------------------------------------------------------- + // #10 — the title + // ------------------------------------------------------------------------- + + describe("readEngineSessionTitle (#10)", () => { + // Table-driven on the VALUE. The bridge is a pass-through: it reports + // exactly what the engine answered and does not decide what "no title" + // means. Collapsing `""` / undefined to `null` is `handleAcpSessionTitle`'s + // `title || null`, pinned in test/routes/session-reads.check.mjs — if the + // bridge started normalising too, one of the two layers would own a rule + // the other also owns, and an "improvement" to one would silently change + // the wire. + const CASES = [ + ["a titled session answers the title verbatim", "mvs_1", "My Title", "My Title"], + ["an untitled session passes null through", "mvs_2", null, null], + ["an empty title passes \"\" through (the route collapses it)", "mvs_3", "", ""], + ["an undefined title passes through undefined", "mvs_4", undefined, undefined], + ]; + + for (const [name, sessionId, engineAnswer, expected] of CASES) { + test(name, async () => { + acpMock.getMcodeSessionTitle = async () => engineAnswer; + const result = await readEngineSessionTitle({ sessionId }); + assert.equal(result.sessionId, sessionId); + assert.equal(result.title, expected); + assert.equal(result.gate.endpoint, "GET /api/acp-session-title"); + }); + } + + test("a missing sessionId never reaches the client", async () => { + let called = 0; + acpMock.getMcodeSessionTitle = async () => { + called++; + return "should not happen"; + }; + const result = await readEngineSessionTitle({ sessionId: "" }); + assert.equal(result.title, null); + assert.equal(called, 0, "an empty id is answered locally, not by a lookup"); + }); + }); + + // ------------------------------------------------------------------------- + // #75 — the version + // ------------------------------------------------------------------------- + + describe("readEngineVersion (#75)", () => { + // Table-driven: [name, agentInfo, expectedVersion]. `unknown` is the + // documented sentinel for "nothing has attached yet" and must stay a + // string — /api/protocol/capabilities uses the same value for the same + // fact, and a monitor semver-parsing the field would throw on null. + const CASES = [ + ["an attached client answers its version", { name: "mcode", version: "0.5.7" }, "0.5.7"], + ["a client without a version answers unknown", { name: "mcode" }, "unknown"], + ["no client at all answers unknown", null, "unknown"], + ]; + + for (const [name, info, expected] of CASES) { + test(name, async () => { + acpMock.getMcodeServerInfo = () => info; + const result = await readEngineVersion(); + assert.equal(result.version, expected); + assert.equal(typeof result.version, "string"); + // The version is a protocol fact. The catalogue host exposes no + // version accessor, so the facade reports the mirror it used + // rather than claiming the engine answered. + assert.equal(result.source, "acp"); + assert.equal(result.gate.gate, "no-capability-key"); + }); + } + }); + + // ------------------------------------------------------------------------- + // host === null — the degradation this batch must not hide + // ------------------------------------------------------------------------- + + describe("catalogue host returned null", () => { + test("a null host is reported as a fallback, never as the engine answering", async () => { + acpMock.getCatalogueHost = async () => null; + acpMock.listAllMcodeSessions = async () => [WIRE_SESSION]; + const result = await readEngineSessionList(); + if (result.transport === "runtime") { + // The transport ASKED for the host and did not get one. The read + // still succeeds from the ACP mirror (the sidebar must not break), + // but the source says so — silently claiming "catalogue" here is + // the fake-success shape. + assert.equal(result.source, "acp-fallback"); + } else { + assert.equal(result.source, "acp"); + } + // Either way the sessions still come back: a read family answers. + assert.deepEqual(result.sessions, [WIRE_SESSION]); + }); + + test("a live host is reported as catalogue under the runtime transport", async () => { + acpMock.getCatalogueHost = async () => ({ adapter: {} }); + acpMock.getMcodeSessionsForWorkspace = async () => [WIRE_SESSION]; + const result = await readEngineSessionListForWorkspace({ cwd: "/ws/a" }); + assert.equal( + result.source, + result.transport === "runtime" ? "catalogue" : "acp", + "the source must follow the transport, not a fixed string", + ); + }); + + test("a host that boots but whose list throws still surfaces the throw", async () => { + acpMock.getCatalogueHost = async () => ({ adapter: {} }); + acpMock.listAllMcodeSessions = async () => { + throw new Error("sqlite locked"); + }; + // The facade does not swallow a broken engine into an empty list — + // that is the #110 fake-success shape. acp-client.js owns the + // ACP failover; the facade adds no second, quieter one. + await assert.rejects(() => readEngineSessionList(), /sqlite locked/); + }); + }); +}); diff --git a/packages/webui/test/lib/engine/session-tree-reads.test.js b/packages/webui/test/lib/engine/session-tree-reads.test.js new file mode 100644 index 00000000..6e84ffe1 --- /dev/null +++ b/packages/webui/test/lib/engine/session-tree-reads.test.js @@ -0,0 +1,881 @@ +// webui/test/lib/engine/session-tree-reads.test.js +// +// M3-B2: the session-tree family's engine facade (GET /api/session-tree). +// +// #8 is the main↔subagent communication spine. The hierarchy the sidebar +// renders is built from `parent_session_id`, so THREE things are pinned +// here, each of them something a refactor could plausibly break while +// looking like a no-op: +// +// 1. The NODE SHAPE. The wire node is exactly +// `{id, title, agent, kind, status, updatedAt, children}` — and in +// particular it carries NO `parent_session_id` key. The hierarchy is +// structural (via `children`), not a field on the node. The batch +// brief asked whether such a key is omitted or `null`; the truthful +// answer, measured against the real 299-node tree before the +// refactor, is that the key does not exist at all. The key SET is +// asserted exactly, not by subset, so both halves stay honest. +// 2. The HIERARCHY FILTER. `buildTree` attaches a child to the root +// session named by its `parent_session_id`, in the SAME directory. +// Anything that does not attach — an orphan (parent not in the row +// set), a cross-directory parent, a grandchild whose parent is +// itself a child, a child of a `root` container row, anything in a +// cycle — is DROPPED SILENTLY. That is long-standing behaviour this +// batch must not change, so it is pinned rather than left to a diff. +// 3. The GATE IS REAL. #8 is 100% engine data, so a provider that +// declares no session listing must produce +// EngineCapabilityNotSupportedError → 501, never an empty tree +// (#110 fake-success). And the route must PROPAGATE that error +// rather than folding it into its own `{ok:false}` soft-fail body — +// that propagation is the one place this batch could have turned a +// 501 into a 200, so it has its own test. +// +// Boundaries probed empirically against the PRE-refactor route, not +// assumed from the batch plan (which was wrong on this point): #8 reads +// exactly one query parameter, `refresh`. `limit`, `offset`, `page` and +// `cursor` are NOT read — `?limit=1` returns the whole tree. The only +// cap is the internal MAX_ROWS = 5000 with `truncated: true`. The +// "limit 缺省/0/超上限" cases below therefore assert the real contract: +// unknown parameters are ignored and nothing truncates below MAX_ROWS. +// +// Test style follows test/lib/engine/session-reads.test.js (batch B1): +// table-driven, one row per case. +// +// Two module-mock traps, both learned in B3 while adding the sibling +// `usage-reads.test.js`, and both recorded here because this suite is where +// a future batch will look for the answer: +// +// 1. `t.mock.module` REPLACES THE WHOLE NAMESPACE, it does not merge. A +// mock that names only the export the test cares about leaves every +// other name undefined, and a consumer that imports more than one name +// from the mocked module then fails at INSTANTIATION with +// `SyntaxError: The requested module '…' does not provide an export +// named '…'` — a failure that reads like a product bug and is not +// one. In this suite it does not bite, because `routes/sessions.js` +// imports exactly one name from `engine/session-tree-reads.js`; in +// `routes/usage.js` it does, because that route binds three reads at +// module scope. When a facade grows a second call, the mock has to +// grow with it — stub the rest with something that throws, so an +// unexpected call is loud instead of returning a plausible payload. +// 2. `mock.module` only re-evaluates the MOCKED specifier. A consumer +// already in the registry keeps its old LIVE BINDING, so a second test +// in the same file silently reuses the first test's mock and passes for +// the wrong reason. Every route re-import below therefore carries a +// fresh `?bust=N`; deleting that query turns nine tests in this file +// red, which is the cheapest proof the mechanism is load-bearing. + +import { test, describe, after } from "node:test"; +import assert from "node:assert/strict"; +import { join } from "node:path"; +import { spawnSync } from "node:child_process"; +import { writeFileSync } from "node:fs"; + +import { mkTmpDir, rmTmpDir } from "../../helpers/tmp.js"; +import { setupMocks, absPath } from "../../helpers/_setup.js"; + +// --------------------------------------------------------------------------- +// Fixture — built BEFORE any server module is imported, and that ordering is +// load-bearing, not stylistic. +// +// `lib/config.js` resolves MCODE_RUNTIME_DB and SESSIONS_DB at MODULE LOAD, +// and `lib/session-tree.js` / `lib/sessions.js` import it statically. So a +// `before()` hook that set the env would be too late: the first static +// import of anything that reaches config.js would already have frozen the +// real ~/.minimax paths, and the fixture would silently read the +// developer's real database. Hence: build the tmp dir, create the db and +// set the env here at module top level, and only then import server code. +// --------------------------------------------------------------------------- + +const tmpDir = mkTmpDir("webui-tree-facade-"); +const dbPath = join(tmpDir, "runtime-state.sqlite"); +const sessionsPath = join(tmpDir, "sessions.json"); + +const FIXTURE_ROWS = [ + // id, parent, type, agent, title, dir-suffix + ["m1", null, "branch", "main-agent", "主会话", ""], + ["c1", "m1", null, "coder", "子 agent", ""], + ["c2", "m1", null, "planner", "子 agent 二", ""], + // a grandchild — must NOT render + ["g1", "c1", null, "tester", "孙 agent", ""], + // an orphan — parent not in the row set + ["orphan", "does-not-exist", null, null, "孤儿", ""], + // a `root` container row and a child hanging off it — neither renders + ["rc", null, "root", null, "容器", ""], + ["under-rc", "rc", null, null, "容器下的子节点", ""], + // a cross-directory child — must NOT render + ["xd", "m1", null, null, "跨目录子节点", "-other"], + // a self-parent and a two-node cycle — neither renders, neither hangs + ["self", "self", null, null, "自环", ""], + ["cyc-a", "cyc-b", null, null, "环 A", ""], + ["cyc-b", "cyc-a", null, null, "环 B", ""], + // title boundaries + ["t-quote", null, "branch", null, 'a"b\\c', ""], + ["t-html", null, "branch", null, "&", ""], + ["t-multi", null, "branch", null, "第一行\n第二行\r\n第三行\t制表", ""], + ["t-emoji", null, "branch", null, "🚀 עברית مرحبا", ""], + ["t-long", null, "branch", null, "长".repeat(5000), ""], + ["t-empty", null, "branch", null, "", ""], + ["t-null", null, "branch", null, null, ""], + // filtered by the SQL WHERE clause — must never reach buildTree + ["f-archived", null, "branch", null, "已归档", ""], + ["f-hidden", null, "branch", null, "不可见", ""], + ["f-peek", null, "branch", null, "peek", ""], + ["f-cron", null, "branch", null, "cron", ""], + ["f-nodir", null, "branch", null, "无目录", ""], +]; + +// Build the db with spawnSync(SQLITE3_BIN) rather than requiring a native +// binding — the same approach test/lib/mcode-session-delete.test.js uses, +// so this suite does not depend on a compiled module being present. +const sqlRows = FIXTURE_ROWS.map(([id, parent, type, agent, title, dirSuffix]) => { + const dir = `${tmpDir}/proj${dirSuffix}`; + const q = (v) => (v === null ? "NULL" : `'${String(v).replace(/'/g, "''")}'`); + return `INSERT INTO local_runtime_sessions + (session_id, record_json, updated_at_ms, agent_name, session_type, status, + archived, visibility, session_kind, parent_session_id, workspace_dir, title, created_at_ms) + VALUES (${q(id)}, '{}', 1000, ${q(agent)}, ${q(type)}, 'idle', + ${id === "f-archived" ? 1 : 0}, + ${id === "f-hidden" ? "'hidden'" : "'visible'"}, + ${id === "f-peek" ? "'peek'" : id === "f-cron" ? "'cron'" : "'conversation'"}, + ${q(parent)}, ${id === "f-nodir" ? "NULL" : q(dir)}, ${q(title)}, 1000);`; +}).join("\n"); + +// sqlite3 has no way to create a file with a schema in one -cmd batch on +// every platform, so the DDL is passed as a single argument like the other +// suites do. +const SQLITE3_BIN = process.env.SQLITE3_BIN || "sqlite3"; +const DDL = ` + CREATE TABLE local_runtime_sessions ( + session_id TEXT PRIMARY KEY, record_json TEXT NOT NULL, + updated_at_ms INTEGER NOT NULL, agent_name TEXT, session_type TEXT, + status TEXT, archived INTEGER NOT NULL DEFAULT 0, + visibility TEXT NOT NULL DEFAULT 'visible', + session_kind TEXT NOT NULL DEFAULT 'conversation', + parent_session_id TEXT, workspace_dir TEXT, title TEXT, created_at_ms INTEGER + ); + ${sqlRows} +`; +{ + const r = spawnSync(SQLITE3_BIN, [dbPath, DDL], { encoding: "utf8" }); + assert.equal(r.status, 0, `sqlite3 create failed: ${r.stderr}`); +} +// A custom title for m1, so the customTitles overlay is exercised too. +writeFileSync( + sessionsPath, + JSON.stringify([ + { id: "w1", mcodeSessionId: "m1", title: "改过名的主会话", titleCustom: true }, + { id: "w2", mcodeSessionId: "c1", title: "不该生效", titleCustom: false }, + ]), +); + +process.env.MCODE_RUNTIME_DB = dbPath; +process.env.MCODE_WEBUI_SESSIONS_DB = sessionsPath; + +// --- now, and only now, the server modules ------------------------------- +const { ENGINE_CAPABILITY_KEYS } = await import("../../../server/engine/index.js"); +const { + SESSION_TREE_ENDPOINTS, + assertSessionTreeCapability, + readEngineSessionTree, + resolveSessionTreeProvider, +} = await import("../../../server/engine/session-tree-reads.js"); +const { buildTree } = await import("../../../server/lib/session-tree.js"); +const { + EngineCapabilityNotSupportedError, + isEngineCapabilityNotSupportedError, + engineCapabilityHttpResponse, +} = await import("../../../server/engine/errors.js"); +const { assertEngineCapability } = await import("../../../server/engine/capabilities.js"); + +const ENDPOINT = "GET /api/session-tree"; + +after(() => { + rmTmpDir(tmpDir); + delete process.env.MCODE_RUNTIME_DB; + delete process.env.MCODE_WEBUI_SESSIONS_DB; +}); + +/** A `buildTree` row, defaulted so each case only states what it is about. */ +const row = (o) => ({ + session_id: o.id, + title: o.title ?? null, + agent_name: o.agent ?? null, + session_kind: o.kind ?? "conversation", + session_type: o.type ?? "branch", + parent_session_id: o.parent ?? null, + workspace_dir: o.dir ?? "/w/proj", + status: o.status ?? "idle", + updated_at_ms: o.at ?? 1, + created_at_ms: o.at ?? 1, +}); + +/** Flatten an assembled tree into `{id, depth, node}` records. */ +function flatten(tree) { + const out = []; + for (const project of tree) { + for (const dir of project.directories) { + for (const session of dir.sessions) { + const walk = (node, depth) => { + out.push({ id: node.id, depth, node }); + for (const child of node.children || []) walk(child, depth + 1); + }; + walk(session, 0); + } + } + } + return out; +} + +/** The session ids the client actually receives, in render order, with depth. */ +const visible = (tree) => flatten(tree).map((n) => [n.id, n.depth]); + +// --------------------------------------------------------------------------- +// 1. The endpoint → capability declaration table +// --------------------------------------------------------------------------- + +describe("SESSION_TREE_ENDPOINTS — this batch's declaration table", () => { + // Table-driven. Editing a row is a capability decision and must be + // reviewed as one, so the table IS the assertion. + const TABLE = [[ENDPOINT, "sessionCrud", "listSessions"]]; + + for (const [endpoint, capability, subItem] of TABLE) { + test(`${endpoint} declares ${capability}.${subItem}`, () => { + const need = SESSION_TREE_ENDPOINTS[endpoint]; + assert.equal(need.capability, capability); + assert.equal(need.subItem, subItem); + }); + } + + test("the table carries exactly the endpoints this batch routes", () => { + assert.deepEqual(Object.keys(SESSION_TREE_ENDPOINTS).sort(), [ENDPOINT]); + }); + + test("the capability is a real key of the 14-key registry", () => { + assert.ok(ENGINE_CAPABILITY_KEYS.includes(SESSION_TREE_ENDPOINTS[ENDPOINT].capability)); + }); +}); + +// --------------------------------------------------------------------------- +// 2. Provider resolution + the hard gate +// --------------------------------------------------------------------------- + +describe("resolveSessionTreeProvider / assertSessionTreeCapability", () => { + // Table-driven. Absent means "no provider claims this transport yet" + // (M4), which is NOT the same answer as "capability unavailable". + const TRANSPORTS = [ + ["runtime", true, "checked"], + ["acp", false, "unregistered-transport"], + ["exec", false, "unregistered-transport"], + ["", false, "unregistered-transport"], + ]; + + for (const [transport, hasProvider, gate] of TRANSPORTS) { + test(`transport "${transport}" → provider=${hasProvider} gate=${gate}`, () => { + const provider = resolveSessionTreeProvider(transport); + assert.equal(provider !== null, hasProvider); + const g = assertSessionTreeCapability(ENDPOINT, transport); + assert.equal(g.gate, gate); + assert.equal(g.endpoint, ENDPOINT); + assert.equal(g.capability, "sessionCrud"); + assert.equal(g.subItem, "listSessions"); + }); + } + + test("an unknown endpoint is caller confusion, not an engine limitation", () => { + // A plain Error, so the HTTP layer never answers 501 for a typo in + // webui's own code. + assert.throws( + () => assertSessionTreeCapability("GET /api/nope", "runtime"), + (err) => { + assert.ok(!(err instanceof EngineCapabilityNotSupportedError)); + assert.equal(err.code, "unknown_session_tree_endpoint"); + assert.match(err.message, /not part of the session-tree family/); + return true; + }, + ); + }); +}); + +describe("the tree gate refuses a provider that cannot list sessions", () => { + // The registered providers declare `full` today, so — exactly as in B1 — + // only this file can prove the gate WOULD bite. A route that answered + // `{ok:true, projects:[]}` would be the #110 failure mode. + const allFull = () => Object.fromEntries(ENGINE_CAPABILITY_KEYS.map((k) => [k, { level: "full" }])); + const PARTIAL_NO_LIST = { + ...allFull(), + sessionCrud: { + level: "partial", + missing: ["listSessions"], + reason: "test fixture: provider exposes no session listing", + }, + }; + const NONE = { + ...allFull(), + sessionCrud: { level: "none", reason: "test fixture: interface-absent" }, + }; + + test("a `none` declaration throws, and maps to 501", () => { + const need = SESSION_TREE_ENDPOINTS[ENDPOINT]; + assert.throws( + () => assertEngineCapability(NONE, need.capability, "fixture-provider"), + (err) => { + assert.ok(isEngineCapabilityNotSupportedError(err)); + assert.equal(err.capability, "sessionCrud"); + assert.equal(err.provider, "fixture-provider"); + const { status, payload } = engineCapabilityHttpResponse(err); + assert.equal(status, 501); + assert.equal(payload.code, "engine_capability_not_supported"); + return true; + }, + ); + }); + + test("a `partial` declaration missing listSessions throws, naming the method", () => { + const need = SESSION_TREE_ENDPOINTS[ENDPOINT]; + assert.throws( + () => assertEngineCapability(PARTIAL_NO_LIST, need.capability, "fixture-provider", need.subItem), + (err) => { + assert.deepEqual(err.missing, ["listSessions"]); + return true; + }, + ); + }); + + test("a `partial` declaration that KEEPS listSessions lets the tree through", () => { + const need = SESSION_TREE_ENDPOINTS[ENDPOINT]; + assert.doesNotThrow(() => + assertEngineCapability( + { ...allFull(), sessionCrud: { level: "partial", missing: ["deleteSession"], reason: "x" } }, + need.capability, + "fixture-provider", + need.subItem, + ), + ); + }); +}); + +// --------------------------------------------------------------------------- +// 3. The red line: node shape, and which rows reach the client +// --------------------------------------------------------------------------- + +describe("buildTree — the node shape the sidebar depends on", () => { + const roots = new Map([["/w/proj", "/w/proj"]]); + + test("a ROOT node carries 7 keys including children, and NO parent_session_id", () => { + const tree = buildTree( + [row({ id: "m1", title: "主会话" }), row({ id: "c1", parent: "m1", title: "子 agent" })], + roots, + ); + const root = flatten(tree)[0].node; + assert.deepEqual( + Object.keys(root).sort(), + ["agent", "children", "id", "kind", "status", "title", "updatedAt"], + ); + assert.equal( + Object.prototype.hasOwnProperty.call(root, "parent_session_id"), + false, + "the node must NOT carry parent_session_id — the tree is structural", + ); + }); + + test("a CHILD node carries 6 keys — it has NO `children` key at all", () => { + // Measured against the real 299-node tree before the refactor: 233 + // root nodes carry `children`, all 66 child nodes do NOT. + // `buildTree` adds `children` only in the output map that wraps each + // ROOT session; a child is pushed into `directory.children` bare and + // never re-wrapped. "Normalising" this — giving every node a + // `children` array — would change 66 nodes' shape in the sidebar, so + // it is pinned here rather than left to a diff. + const tree = buildTree([row({ id: "m1" }), row({ id: "c1", parent: "m1" })], roots); + const child = flatten(tree).find((n) => n.depth === 1).node; + assert.deepEqual( + Object.keys(child).sort(), + ["agent", "id", "kind", "status", "title", "updatedAt"], + ); + assert.equal( + Object.prototype.hasOwnProperty.call(child, "children"), + false, + "a child node must not gain a children key — that is a client-visible shape change", + ); + }); + + test("a leaf ROOT's children is an empty array, not null and not missing", () => { + const leaf = flatten(buildTree([row({ id: "m1" })], roots))[0].node; + assert.deepEqual(leaf.children, []); + }); + + test("a null title becomes an empty string", () => { + const n = flatten(buildTree([row({ id: "m1", title: null })], roots))[0].node; + assert.equal(n.title, ""); + assert.equal(typeof n.title, "string"); + }); +}); + +describe("buildTree — which rows reach the client (the subagent hierarchy)", () => { + const W = "/w/proj"; + const W2 = "/w/other"; + const roots = new Map([ + [W, W], + [W2, W2], + ]); + + // Table-driven over the boundary cases. `expected` is the set of ids the + // client actually receives, at the depth it receives them. A row that + // vanishes is a subagent the user cannot see; a row that arrives at the + // wrong depth is the same defect. Both are pinned. + const CASES = [ + { + name: "a child attaches to its parent at depth 1", + rows: [row({ id: "m1" }), row({ id: "c1", parent: "m1" })], + expected: [["m1", 0], ["c1", 1]], + }, + { + name: "an orphan (parent not in the row set) is dropped", + rows: [row({ id: "m1" }), row({ id: "orphan", parent: "gone" })], + expected: [["m1", 0]], + }, + { + name: "a child whose parent is in ANOTHER directory is dropped", + rows: [row({ id: "m1" }), row({ id: "x", parent: "m1", dir: W2 })], + expected: [["m1", 0]], + }, + { + name: "a grandchild is dropped — only ONE level of subagent renders", + rows: [row({ id: "m1" }), row({ id: "c1", parent: "m1" }), row({ id: "g1", parent: "c1" })], + expected: [["m1", 0], ["c1", 1]], + }, + { + name: "a child of a `root` container row is dropped", + rows: [row({ id: "rc", type: "root" }), row({ id: "u", parent: "rc" })], + expected: [], + }, + { + name: "a `root` container row is itself not a sidebar entry", + rows: [row({ id: "rc", type: "root" }), row({ id: "m1" })], + expected: [["m1", 0]], + }, + { + name: "a self-parenting row is dropped, and does not hang the build", + rows: [row({ id: "m1" }), row({ id: "self", parent: "self" })], + expected: [["m1", 0]], + }, + { + name: "a two-node cycle is dropped and does not hang the build", + rows: [row({ id: "m1" }), row({ id: "a", parent: "b" }), row({ id: "b", parent: "a" })], + expected: [["m1", 0]], + }, + { + name: "several children of one parent all render, sorted by recency", + rows: [ + row({ id: "m1" }), + row({ id: "old", parent: "m1", at: 1 }), + row({ id: "new", parent: "m1", at: 9 }), + ], + expected: [["m1", 0], ["new", 1], ["old", 1]], + }, + ]; + + for (const { name, rows, expected } of CASES) { + test(name, () => { + assert.deepEqual(visible(buildTree(rows, roots)), expected); + }); + } + + test("the depth distribution is exactly two levels for a main+subagent+deeper shape", () => { + // The "层深分布" the batch brief asks to be compared: whatever the db + // holds, the client never sees deeper than depth 1. + const rows = [ + row({ id: "m1" }), + row({ id: "m2" }), + row({ id: "c1", parent: "m1" }), + row({ id: "g1", parent: "c1" }), + row({ id: "g2", parent: "g1" }), + row({ id: "orphan", parent: "nope" }), + ]; + const hist = {}; + for (const n of flatten(buildTree(rows, roots))) { + hist[n.depth] = (hist[n.depth] || 0) + 1; + } + assert.deepEqual(hist, { 0: 2, 1: 1 }); + }); +}); + +describe("buildTree — title and field boundaries", () => { + const roots = new Map([["/w/proj", "/w/proj"]]); + + // Table-driven: [name, title, expected]. A title travels into both export + // formats and into the sidebar label, so special characters, newlines and + // absurd lengths must survive verbatim rather than be normalised. + const TITLES = [ + ["plain", "普通标题", "普通标题"], + ["quotes and backslash", 'a"b\\c', 'a"b\\c'], + ["html-ish", "&", "&"], + ["multi-line", "第一行\n第二行\r\n第三行\t制表", "第一行\n第二行\r\n第三行\t制表"], + ["emoji and rtl", "🚀 עברית مرحبا", "🚀 עברית مرحبا"], + ["§§ marker-looking", "§§ turn_msg=abc", "§§ turn_msg=abc"], + ["very long", "长".repeat(5000), "长".repeat(5000)], + ["empty", "", ""], + ["null becomes empty", null, ""], + ["only whitespace", " ", " "], + ]; + + for (const [name, title, expected] of TITLES) { + test(`title: ${name}`, () => { + const n = flatten(buildTree([row({ id: "m1", title })], roots))[0].node; + assert.equal(n.title, expected); + }); + } + + // A row that states ONLY the columns it must — no defaulting helper, so + // a column really is absent rather than filled in with a placeholder. + // `buildTree` reads `row.agent_name || ""`, `row.session_kind || ""`, + // `row.status || ""` and `row.updated_at_ms ?? 0`, so absent and + // empty-string collapse to the same node value; `updated_at_ms` is the + // one that distinguishes missing (0) from falsy-but-present. + const bare = (o) => ({ + session_id: o.id, + parent_session_id: o.parent ?? null, + session_type: o.type ?? "branch", + workspace_dir: o.dir ?? "/w/proj", + ...o.extra, + }); + + // Table-driven: [name, bare-row, field, expected]. + const FIELDS = [ + ["agent_name absent → empty string", bare({ id: "m1" }), "agent", ""], + ["agent_name empty → empty string", bare({ id: "m1", extra: { agent_name: "" } }), "agent", ""], + ["agent_name present", bare({ id: "m1", extra: { agent_name: "coder" } }), "agent", "coder"], + ["session_kind absent → empty string", bare({ id: "m1" }), "kind", ""], + ["session_kind task", bare({ id: "m1", extra: { session_kind: "task" } }), "kind", "task"], + ["status absent → empty string", bare({ id: "m1" }), "status", ""], + ["status running", bare({ id: "m1", extra: { status: "running" } }), "status", "running"], + ["updated_at_ms absent → 0", bare({ id: "m1" }), "updatedAt", 0], + ["updated_at_ms 0 stays 0", bare({ id: "m1", extra: { updated_at_ms: 0 } }), "updatedAt", 0], + ["title absent → empty string", bare({ id: "m1" }), "title", ""], + ["title null → empty string", bare({ id: "m1", extra: { title: null } }), "title", ""], + ]; + + for (const [name, r, field, expected] of FIELDS) { + test(name, () => { + assert.equal(flatten(buildTree([r], roots))[0].node[field], expected); + }); + } +}); + +describe("buildTree — empty and single-session inputs", () => { + const roots = new Map([["/w/proj", "/w/proj"]]); + + test("no rows at all → an empty project list, not null and not a throw", () => { + assert.deepEqual(buildTree([], roots), []); + }); + + test("a single main session → one project, one directory, one session", () => { + const tree = buildTree([row({ id: "m1" })], roots); + assert.equal(tree.length, 1); + assert.equal(tree[0].directories.length, 1); + assert.equal(tree[0].directories[0].sessions.length, 1); + assert.equal(tree[0].sessionCount, 1); + }); + + test("a single main session with children reports sessionCount 1, not 3", () => { + // The project pill counts user-started sessions; subagents must not + // inflate it. Pinned because it is easy to "fix" by accident. + const tree = buildTree( + [row({ id: "m1" }), row({ id: "c1", parent: "m1" }), row({ id: "c2", parent: "m1" })], + roots, + ); + assert.equal(tree[0].sessionCount, 1); + assert.equal(tree[0].directories[0].sessions[0].children.length, 2); + }); + + test("only orphan rows → no sessions, but the directory still appears", () => { + const tree = buildTree([row({ id: "o1", parent: "gone" })], roots); + assert.equal(tree.length, 1, "the directory is a grouping key even with no visible session"); + assert.deepEqual(tree[0].directories[0].sessions, []); + assert.equal(tree[0].sessionCount, 0); + }); +}); + +// --------------------------------------------------------------------------- +// 4. The facade: forwarding, verbatim, against a real db +// --------------------------------------------------------------------------- + +describe("readEngineSessionTree — forwards the payload verbatim", () => { + test("the result carries the tree plus source/gate/transport", async () => { + const { tree, source, gate, transport } = await readEngineSessionTree({ force: true }); + assert.equal(tree.ok, true); + assert.equal(source, "runtime-db", "the tree is not a transport-switched surface"); + assert.equal(gate.endpoint, ENDPOINT); + assert.equal(typeof transport, "string"); + assert.deepEqual(Object.keys(tree).sort(), [ + "cached", + "counts", + "generatedAt", + "ok", + "projects", + "truncated", + ]); + }); + + test("the forwarded tree is the 1-level shape buildTree produces", async () => { + // The fixture db carries 25 seeded rows; the SQL WHERE clause drops the + // 5 filtered ones, and the hierarchy filter drops the orphan, the + // grandchild, the cross-directory child, the two cycle rows, the + // self-parent, the `root` container and its child. Only main sessions + // and their direct children may appear. + const { tree } = await readEngineSessionTree({ force: true }); + const ids = visible(tree.projects).map(([id]) => id); + assert.ok(!ids.includes("orphan"), "an orphan must not reach the client"); + assert.ok(!ids.includes("g1"), "a grandchild must not reach the client"); + assert.ok(!ids.includes("xd"), "a cross-directory child must not reach the client"); + assert.ok(!ids.includes("cyc-a") && !ids.includes("cyc-b"), "cycle rows must not reach the client"); + assert.ok(!ids.includes("rc") && !ids.includes("under-rc"), "container rows must not reach the client"); + assert.ok(!ids.includes("f-archived"), "archived rows are filtered by SQL"); + assert.ok(!ids.includes("f-peek"), "peek rows are filtered by SQL"); + assert.ok(!ids.includes("f-nodir"), "rows without a directory are filtered by SQL"); + // Exactly two depths, and the children hang off m1. + const depths = new Set(visible(tree.projects).map(([, d]) => d)); + assert.deepEqual([...depths].sort(), [0, 1]); + const m1 = flatten(tree.projects).find((n) => n.id === "m1"); + assert.deepEqual(m1.node.children.map((c) => c.id).sort(), ["c1", "c2"]); + assert.equal(tree.truncated, false, "a small fixture never truncates"); + }); + + test("a custom title from the webui store overlays the db title", async () => { + // `customTitles` only exists on the webui record, so without the + // overlay the sidebar would keep showing whatever mcode generated. + const { tree } = await readEngineSessionTree({ force: true }); + const m1 = flatten(tree.projects).find((n) => n.id === "m1"); + assert.equal(m1.node.title, "改过名的主会话"); + const c1 = flatten(tree.projects).find((n) => n.id === "c1"); + assert.equal(c1.node.title, "子 agent", "titleCustom:false must NOT overlay"); + }); + + test("force:false reuses the 15s cache and reports cached:true", async () => { + const first = await readEngineSessionTree({ force: true }); + assert.equal(first.tree.cached, false); + const second = await readEngineSessionTree({ force: false }); + assert.equal(second.tree.cached, true, "the facade forwards the cache flag, it does not bypass the cache"); + }); +}); + +// --------------------------------------------------------------------------- +// 5. The route: pass-through, and the 501 that must NOT be swallowed +// --------------------------------------------------------------------------- + +describe("handleSessionTree — the route passes the facade payload through", () => { + // One fresh route module per test. node:test's `mock.module` re-evaluates + // the MOCKED specifier, but a route module already sitting in the registry + // keeps its old live binding to the facade — so the second and third tests + // in this suite would silently exercise the FIRST test's mock and pass for + // the wrong reason. The `?bust=N` query makes the route re-resolve the + // facade specifier, which is what picks up the new mock. (These tests need + // the `--experimental-test-module-mocks` flag that the `test:unit` and + // `test` scripts already pass.) + // + // A mutation that deletes the `?bust=N` from this suite's re-import turns + // NINE of its tests red at once; that is the cheapest proof the mechanism + // is load-bearing rather than decorative. + let bust = 0; + + test("a facade payload is written to the response byte-for-byte", async (t) => { + // The payload is injected rather than produced, so this is about the + // ROUTE's contract: it must not re-shape, re-count or re-derive + // anything. The payload carries the exact key set the real tree + // produces, `cached` included. + const payload = { + ok: true, + generatedAt: 1750000000000, + truncated: false, + counts: { projects: 1, directories: 1, sessions: 2 }, + projects: [ + { + key: "proj", + name: "proj", + repoPaths: ["/w/proj"], + latestAt: 1000, + sessionCount: 1, + directories: [ + { + path: "/w/proj", + name: "proj", + latestAt: 1000, + sessions: [ + { + id: "m1", + title: "主会话", + agent: "", + kind: "conversation", + status: "idle", + updatedAt: 1000, + children: [ + { + id: "c1", + title: "子 agent", + agent: "coder", + kind: "task", + status: "running", + updatedAt: 900, + children: [], + }, + ], + }, + ], + }, + ], + }, + ], + cached: false, + }; + await setupMocks(t, { acp: {} }); + t.mock.module(absPath("engine/session-tree-reads.js"), { + namedExports: { + readEngineSessionTree: async () => ({ + tree: payload, + source: "runtime-db", + gate: { endpoint: ENDPOINT, gate: "checked" }, + transport: "acp", + }), + }, + }); + const sessionsRoute = await import(`${absPath("routes/sessions.js")}?bust=${bust++}`); + const written = []; + const res = { + headersSent: false, + writeHead(status, headers) { written.push({ status, headers }); this.headersSent = true; return this; }, + end(body) { written.push({ body }); return this; }, + }; + await sessionsRoute.handleSessionTree({ url: "/api/session-tree" }, res, { cid: "t" }); + assert.equal(written[0].status, 200); + assert.equal(written[0].headers["Cache-Control"], "no-store"); + assert.deepEqual(JSON.parse(written[1].body), payload); + }); + + test("?refresh=1 reaches the facade as force:true, and nothing else does", async (t) => { + await setupMocks(t, { acp: {} }); + const seen = []; + t.mock.module(absPath("engine/session-tree-reads.js"), { + namedExports: { + readEngineSessionTree: async (o) => { + seen.push(o); + return { tree: { ok: true }, source: "runtime-db", gate: {}, transport: "acp" }; + }, + }, + }); + const sessionsRoute = await import(`${absPath("routes/sessions.js")}?bust=${bust++}`); + const mk = () => ({ + headersSent: false, + writeHead() { this.headersSent = true; return this; }, + end() { return this; }, + }); + // Table-driven: [query, expected force]. The limit/offset/page/cursor + // rows are the measured contract, not an assumption — #8 never read + // them, and adding a clamp here would invent behaviour. + const QUERIES = [ + ["?refresh=1", true], + ["", false], + ["?refresh=0", false], + ["?refresh=true", false], + ["?limit=1", false], + ["?limit=0", false], + ["?limit=999999", false], + ["?offset=5", false], + ["?page=2", false], + ["?cursor=x", false], + ["?limit=1&refresh=1", true], + ]; + for (const [q] of QUERIES) { + await sessionsRoute.handleSessionTree({ url: `/api/session-tree${q}` }, mk(), { cid: "t" }); + } + assert.equal(seen.length, QUERIES.length); + for (let i = 0; i < QUERIES.length; i += 1) { + assert.equal(seen[i].force, QUERIES[i][1], `query "${QUERIES[i][0]}" → force=${QUERIES[i][1]}`); + } + }); + + test("a capability error PROPAGATES so invokeHandler can answer 501", async (t) => { + // The one place this batch could have turned a 501 into a 200: the + // route's try/catch would fold the gate error into its own + // `{ok:false, reason:"session_tree_failed"}` body. It must not — the + // declaration gate is the whole point of the batch. + await setupMocks(t, { acp: {} }); + t.mock.module(absPath("engine/session-tree-reads.js"), { + namedExports: { + readEngineSessionTree: async () => { + throw new EngineCapabilityNotSupportedError({ + capability: "sessionCrud", + provider: "fixture-provider", + missing: ["listSessions"], + reason: "test fixture: interface-absent", + }); + }, + }, + }); + const sessionsRoute = await import(`${absPath("routes/sessions.js")}?bust=${bust++}`); + const res = { + headersSent: false, + writeHead() { this.headersSent = true; return this; }, + end() { return this; }, + }; + await assert.rejects( + () => sessionsRoute.handleSessionTree({ url: "/api/session-tree" }, res, { cid: "t" }), + isEngineCapabilityNotSupportedError, + ); + }); + + test("a LOOKALIKE error that merely carries the right .name does NOT propagate", async (t) => { + // The route discriminates with `isEngineCapabilityNotSupportedError` + // (an `instanceof` check), not with `cause.name === "…"`. `.name` is a + // writable instance property, so any code upstream can make an ordinary + // error impersonate the gate's — and a `.name` compare would then + // re-throw it and turn a soft-fail into a 501 the engine never + // declared. Pinned as a pair with the test above: the real class + // propagates, the impersonator does not. + await setupMocks(t, { acp: {} }); + const lookalike = new Error("not the gate"); + lookalike.name = "EngineCapabilityNotSupportedError"; + t.mock.module(absPath("engine/session-tree-reads.js"), { + namedExports: { readEngineSessionTree: async () => { throw lookalike; } }, + }); + const sessionsRoute = await import(`${absPath("routes/sessions.js")}?bust=${bust++}`); + const written = []; + const res = { + headersSent: false, + writeHead(s, h) { written.push({ s, h }); this.headersSent = true; return this; }, + end(b) { written.push({ b }); return this; }, + }; + // Must NOT reject: an impostor is an ordinary failure and degrades. + await sessionsRoute.handleSessionTree({ url: "/api/session-tree" }, res, { cid: "t" }); + assert.equal(written[0].s, 200, "an impostor must not become a 501"); + const body = JSON.parse(written[1].b); + assert.equal(body.ok, false); + assert.equal(body.reason, "session_tree_failed"); + assert.equal(body.detail, "not the gate"); + }); + + test("a NON-capability failure still degrades to ok:false + reason", async (t) => { + // The soft-fail contract for a broken tree read is unchanged: 200 with + // `{ok:false, reason:"session_tree_failed"}`. + await setupMocks(t, { acp: {} }); + t.mock.module(absPath("engine/session-tree-reads.js"), { + namedExports: { + readEngineSessionTree: async () => { + throw new Error("boom"); + }, + }, + }); + const sessionsRoute = await import(`${absPath("routes/sessions.js")}?bust=${bust++}`); + const written = []; + const res = { + headersSent: false, + writeHead(s, h) { written.push({ s, h }); this.headersSent = true; return this; }, + end(b) { written.push({ b }); return this; }, + }; + await sessionsRoute.handleSessionTree({ url: "/api/session-tree" }, res, { cid: "t" }); + assert.equal(written[0].s, 200); + const body = JSON.parse(written[1].b); + assert.equal(body.ok, false); + assert.equal(body.reason, "session_tree_failed"); + assert.equal(body.detail, "boom"); + }); +}); diff --git a/packages/webui/test/lib/engine/usage-reads.test.js b/packages/webui/test/lib/engine/usage-reads.test.js new file mode 100644 index 00000000..9572031f --- /dev/null +++ b/packages/webui/test/lib/engine/usage-reads.test.js @@ -0,0 +1,1230 @@ +// webui/test/lib/engine/usage-reads.test.js +// +// M3-B3: the usage family's engine facade (#15, #16, #17, #19). +// +// This family is the batch where a "harmless" refactor can be entirely +// silent, because three of its four numbers are DERIVED and none of them +// is compared against anything. So the four things pinned here are: +// +// 1. THE FORMULA'S INPUTS. `contextUsed` is `totalInput + totalOutput + +// totalReasoning` — the CUMULATIVE figure, not the chat flow's +// per-turn `lastTurnContextTokens`, and explicitly NOT including the +// cache counters (which are a subset of `input` and would +// double-count). Section 3 does not assert the formula's result for a +// handful of inputs; it perturbs each of the seven numeric fields one +// at a time and records WHICH ones move the answer. A future +// "simplification" that swaps in the per-turn figure, or that starts +// adding `totalCacheRead`, cannot pass. +// +// 2. THE NUMERIC SNAPSHOT on a real sqlite fixture, row by row, for the +// boundary cases the endpoint exists for: reasoning present, reasoning +// absent, cache hit zero, single turn, many turns, a session whose +// row is gone but whose usage rows remain, and NULL token columns. +// The expected values are written out longhand, not recomputed by the +// same expression under test — a test that computes its oracle with +// the implementation's formula proves nothing. +// +// 3. THE FORECAST SEQUENCE. #19 is a pure function of a history prefix, +// so consecutive reads of a growing history must move the way the +// pre-refactor implementation moved them: no re-filtering, no +// re-sorting, no re-sampling. Section 5 walks every prefix and +// compares against the module's own `forecastExhaustion(readHistory())`. +// +// 4. THE GATE IS REAL, AND THE MOCK IS REAL. The registered provider +// declares `usageStats` and `authCredentials` `full`, so only this +// file can prove the gate would bite. And node:test's +// `mock.module` re-evaluates only the MOCKED specifier, so a route +// module already in the registry keeps its old live binding — every +// route test here re-imports the route under a fresh `?bust=N`, and +// section 6 ends with the control that proves the mock took: with no +// mock at all, the same request reads the fixture db. +// +// Test style follows test/lib/engine/session-reads.test.js (B1) and +// test/lib/engine/session-tree-reads.test.js (B2): table-driven, one row +// per case, fixture built before any server module is imported. + +import { test, describe, after } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync } from "node:fs"; +import { join } from "node:path"; +import { Readable } from "node:stream"; +import { DatabaseSync } from "node:sqlite"; + +import { mkTmpDir, rmTmpDir } from "../../helpers/tmp.js"; +import { setupMocks, absPath } from "../../helpers/_setup.js"; + +// --------------------------------------------------------------------------- +// Fixture — built BEFORE any server module is imported, and that ordering is +// load-bearing, not stylistic. +// +// `lib/config.js` resolves MAVIS_DB_PATH at MODULE LOAD from +// `MINIMAX_DATA_DIR ?? MAVIS_DATA_DIR`, and `lib/mavis-usage.js` imports it +// statically. A `before()` hook that set the env would be too late: the +// first import reaching config.js would already have frozen the real +// ~/.minimax path, and every case below would read the developer's own +// database instead of the fixture. Hence: build the dir and the db, set the +// env here at module top level, and only then import server code. +// +// BOTH env names are set, not just MAVIS_DATA_DIR — `MINIMAX_DATA_DIR` +// wins, and a gate command that isolates the runtime data dir exports it. +// A fixture that wants the database owns the variable that wins. +// +// Prefixes are registered in scripts/test-tmp-leak.check.mjs#KNOWN_PREFIXES; +// a new prefix without that entry fails the test:release-tools gate. +// --------------------------------------------------------------------------- + +const tmpDir = mkTmpDir("mcode-webui-usage-"); +const histDir = mkTmpDir("webui-quota-forecast-test-"); +const dbPath = join(tmpDir, "v2", "sqlite", "runtime-state.sqlite"); +mkdirSync(join(tmpDir, "v2", "sqlite"), { recursive: true }); + +// T0 is a fixed instant, never Date.now(): every expected number below is +// written longhand, and a moving clock would make the fixture unreviewable. +const T0 = 1700000000000; + +/** + * The boundary rows. Every id matches `mvs_[a-f0-9]{16,}` because + * `lib/mavis-usage.js` refuses anything else — the rejection is one of the + * pinned behaviours, not an accident of the fixture. + * + * The token columns are declared NULLABLE on purpose. The shipped v2 schema + * declares them NOT NULL, but `mavis-usage.js` coerces with + * `Number(x) || 0`, so a NULL written by any other writer is a live code + * path; the `null-token-columns` row exercises it through the real query. + */ +const USAGE_ROWS = [ + // [sid, turnId, ts, in, out, reasoning, cacheRead, cacheWrite, model] + // Three turns, all with reasoning. Totals: in 6000, out 2100, reasoning + // 2700, cacheRead 30, cacheWrite 5 → contextUsed 10800. + ["mvs_1111111111111111aaaaaaaaaaaaaa1", "t1", T0, 1000, 500, 300, 0, 0, "MiniMax-M3"], + ["mvs_1111111111111111aaaaaaaaaaaaaa1", "t2", T0 + 1000, 2000, 700, 900, 10, 0, "MiniMax-M3"], + ["mvs_1111111111111111aaaaaaaaaaaaaa1", "t3", T0 + 2000, 3000, 900, 1500, 20, 5, "MiniMax-M3"], + // Two turns where ONLY the first has reasoning. Totals: in 122, out 24, + // reasoning 333, cacheRead 499 → contextUsed 479. The per-turn figure for + // the LAST turn is 11+2+0 = 13, so this row is the one that separates + // "cumulative" from "per turn" by a factor of 36. + ["mvs_2222222222222222bbbbbbbbbbbbbbb2", "t1", T0, 111, 22, 333, 444, 0, "MiniMax-M2.7"], + ["mvs_2222222222222222bbbbbbbbbbbbbbb2", "t2", T0 + 1000, 11, 2, 0, 55, 0, "MiniMax-M2.7"], + // Every counter zero. contextUsed 0, and the 0 must not be confused + // with "no rows" (which is found:false). + ["mvs_3333333333333333ccccccccccccccc3", "t1", T0, 0, 0, 0, 0, 0, "MiniMax-M3"], + // NULL token columns → every total is 0 after `Number(null) || 0`. + ["mvs_4444444444444444ddddddddddddddd4", "t1", T0, null, null, null, null, null, "MiniMax-M3"], + // Usage rows whose session row is GONE (the delete left them behind). + // Totals: in 4242, out 84, reasoning 21, cacheRead 7 → contextUsed 4347. + ["mvs_5555555555555555eeeeeeeeeeeeeee5", "t1", T0, 4242, 84, 21, 7, 0, "MiniMax-M3"], + // Two turns, both with a model, cache never hit. + ["mvs_6666666666666666fffffffffffffff6", "t1", T0, 500, 50, 5, 0, 0, "MiniMax-M2.7-highspeed"], + ["mvs_6666666666666666fffffffffffffff6", "t2", T0 + 1000, 600, 60, 6, 0, 0, "MiniMax-M2.7-highspeed"], + // A single turn with reasoning — the smallest row that still has all three + // summands non-zero. + ["mvs_7777777777777777aaaaaaaaaaaaaaaa7", "t1", T0, 9, 3, 4, 0, 0, "MiniMax-M3"], +]; + +// Sessions that still exist. The orphan's id is deliberately absent. +// Deduped: a session with several usage rows must still be one session row. +const LIVE_SESSIONS = [...new Set(USAGE_ROWS.map((r) => r[0]))].filter( + (sid) => sid !== "mvs_5555555555555555eeeeeeeeeeeeeee5", +); + +{ + const db = new DatabaseSync(dbPath); + db.exec(` + CREATE TABLE local_runtime_token_usage ( + id INTEGER PRIMARY KEY AUTOINCREMENT, + session_id TEXT, agent_name TEXT, framework_type TEXT, turn_id TEXT, + model TEXT, ts INTEGER, input_tokens INTEGER, output_tokens INTEGER, + reasoning_tokens INTEGER, cache_read_tokens INTEGER, + cache_write_tokens INTEGER, cost_usd REAL, raw TEXT + ); + CREATE TABLE local_runtime_sessions (session_id TEXT PRIMARY KEY, title TEXT); + `); + const ins = db.prepare( + `INSERT INTO local_runtime_token_usage + (session_id, agent_name, framework_type, turn_id, model, ts, + input_tokens, output_tokens, reasoning_tokens, cache_read_tokens, cache_write_tokens) + VALUES (?, 'main', 'pi-agent', ?, ?, ?, ?, ?, ?, ?, ?)`, + ); + for (const [sid, turn, ts, i, o, r, cr, cw, model] of USAGE_ROWS) { + ins.run(sid, turn, model, ts, i, o, r, cr, cw); + } + const insS = db.prepare("INSERT INTO local_runtime_sessions (session_id, title) VALUES (?, ?)"); + for (const sid of LIVE_SESSIONS) insS.run(sid, "t"); + db.close(); +} + +process.env.MINIMAX_DATA_DIR = tmpDir; +process.env.MAVIS_DATA_DIR = tmpDir; +process.env.MCODE_WEBUI_HISTORY_PATH = join(histDir, "usage-history.ndjson"); + +// --- now, and only now, the server modules ------------------------------- +const { ENGINE_CAPABILITY_KEYS } = await import("../../../server/engine/index.js"); +const { + USAGE_READ_ENDPOINTS, + assertUsageReadCapability, + contextUsedTokens, + readEngineAccountQuota, + readEngineQuotaForecast, + readEngineSessionUsage, + resolveUsageReadProvider, +} = await import("../../../server/engine/usage-reads.js"); +const { + EngineCapabilityNotSupportedError, + isEngineCapabilityNotSupportedError, + engineCapabilityHttpResponse, +} = await import("../../../server/engine/errors.js"); +const { assertEngineCapability } = await import("../../../server/engine/capabilities.js"); +const { forecastExhaustion, readHistory, appendHistory } = await import( + "../../../server/lib/quota-forecast.js" +); + +const RUNTIME = "runtime"; + +after(() => { + rmTmpDir(tmpDir); + rmTmpDir(histDir); + delete process.env.MINIMAX_DATA_DIR; + delete process.env.MAVIS_DATA_DIR; + delete process.env.MCODE_WEBUI_HISTORY_PATH; +}); + +// --------------------------------------------------------------------------- +// 1. The endpoint → capability declaration table +// --------------------------------------------------------------------------- + +describe("USAGE_READ_ENDPOINTS — this batch's declaration table", () => { + test("covers exactly the four endpoints of batch B3", () => { + assert.deepEqual(Object.keys(USAGE_READ_ENDPOINTS).sort(), [ + "GET /api/usage-real", + "GET /api/usage/forecast", + "POST /api/usage", + "POST /api/usage-trigger", + ]); + }); + + // Table-driven. Editing a row is a capability decision and must be + // reviewed as one, so the table IS the assertion. + const TABLE = [ + ["POST /api/usage", "authCredentials", "getAccountStatus"], + ["POST /api/usage-trigger", "authCredentials", "getAccountStatus"], + ["GET /api/usage-real", "usageStats", "getSessionUsage"], + ]; + for (const [endpoint, capability, subItem] of TABLE) { + test(`${endpoint} declares ${capability}.${subItem}`, () => { + assert.deepEqual(USAGE_READ_ENDPOINTS[endpoint], { capability, subItem }); + // The capability must be one of the 14 matrix keys — the table must + // not grow a private key, which validateEngineCapabilities exists to + // prevent. + assert.ok(ENGINE_CAPABILITY_KEYS.includes(capability)); + }); + } + + test("GET /api/usage/forecast declares NO capability, and the gate says so", () => { + // The forecast reads webui's OWN usage-history.ndjson and calls no + // engine surface. Declaring a capability here would put a lie in the + // registry; gating it hard would remove a working endpoint in response + // to a declaration about something it does not depend on. The value is + // `null`, exactly as B1's `/api/health` — and the gate reports the + // no-op rather than silently passing. + assert.equal(USAGE_READ_ENDPOINTS["GET /api/usage/forecast"], null); + for (const transport of [RUNTIME, "acp", "exec", ""]) { + const g = assertUsageReadCapability("GET /api/usage/forecast", transport); + assert.equal(g.gate, "no-capability-key"); + assert.equal(g.capability, null); + assert.equal(g.subItem, null); + } + }); + + test("an endpoint outside this family is caller confusion, not an engine limitation", () => { + assert.throws( + () => assertUsageReadCapability("GET /api/nope", RUNTIME), + (err) => { + assert.ok(!(err instanceof EngineCapabilityNotSupportedError)); + assert.equal(err.code, "unknown_usage_read_endpoint"); + assert.match(err.message, /not part of the usage family/); + return true; + }, + ); + }); +}); + +// --------------------------------------------------------------------------- +// 2. Provider resolution + the gate +// --------------------------------------------------------------------------- + +describe("resolveUsageReadProvider / assertUsageReadCapability", () => { + // Table-driven. Absent means "no provider claims this transport yet" + // (M4), which is NOT the same answer as "capability unavailable" — the + // default `acp` transport must keep working, so it must NOT throw. + const TRANSPORTS = [ + [RUNTIME, true, "checked"], + ["acp", false, "unregistered-transport"], + ["exec", false, "unregistered-transport"], + ["", false, "unregistered-transport"], + ]; + for (const [transport, hasProvider, gate] of TRANSPORTS) { + test(`transport "${transport}" → provider=${hasProvider} gate=${gate}`, () => { + assert.equal(resolveUsageReadProvider(transport) !== null, hasProvider); + const g = assertUsageReadCapability("GET /api/usage-real", transport); + assert.equal(g.gate, gate); + assert.equal(g.capability, "usageStats"); + assert.equal(g.subItem, "getSessionUsage"); + }); + } +}); + +describe("the usage gate refuses a provider that cannot report usage", () => { + // The registered providers declare `full` today, so — exactly as in B1 and + // B2 — only this file can prove the gate WOULD bite. + const allFull = () => Object.fromEntries(ENGINE_CAPABILITY_KEYS.map((k) => [k, { level: "full" }])); + const withUsage = (usageStats, authCredentials = { level: "full" }) => ({ + ...allFull(), + usageStats, + authCredentials, + }); + + // Table-driven over (endpoint, capability, subItem, declaration). + const CASES = [ + [ + "GET /api/usage-real", + "a `none` usageStats throws and maps to 501", + { level: "none", reason: "test fixture: interface-absent" }, + undefined, + ], + [ + "GET /api/usage-real", + "a `partial` usageStats missing getSessionUsage throws, naming the method", + { level: "partial", missing: ["getSessionUsage"], reason: "test fixture: no per-session usage" }, + "getSessionUsage", + ], + [ + "POST /api/usage", + "a `none` authCredentials throws and maps to 501", + { level: "full" }, + undefined, + { level: "none", reason: "test fixture: interface-absent" }, + ], + [ + "POST /api/usage-trigger", + "a `partial` authCredentials missing getAccountStatus throws", + { level: "full" }, + "getAccountStatus", + { level: "partial", missing: ["getAccountStatus"], reason: "test fixture: no account status" }, + ], + ]; + + for (const [endpoint, title, usageStats, subItem, auth] of CASES) { + test(title, () => { + const need = USAGE_READ_ENDPOINTS[endpoint]; + const decl = withUsage(usageStats, auth || { level: "full" }); + assert.throws( + () => assertEngineCapability(decl, need.capability, "fixture-provider", need.subItem), + (err) => { + assert.ok(isEngineCapabilityNotSupportedError(err), "the real class, so invokeHandler's instanceof matches"); + assert.equal(err.capability, need.capability); + assert.equal(err.provider, "fixture-provider"); + if (subItem) assert.deepEqual(err.missing, [subItem]); + const { status, payload } = engineCapabilityHttpResponse(err); + assert.equal(status, 501); + assert.equal(payload.code, "engine_capability_not_supported"); + return true; + }, + ); + }); + } + + test("a `partial` that KEEPS the sub-item lets the read through", () => { + assert.doesNotThrow(() => + assertEngineCapability( + withUsage( + { level: "partial", missing: ["watchSessionUsageCommits"], reason: "x" }, + { level: "partial", missing: ["listModelProviders"], reason: "y" }, + ), + "usageStats", + "fixture-provider", + "getSessionUsage", + ), + ); + }); + + test("an error that merely carries the right .name is NOT the gate's error", () => { + // `.name` is a writable instance property, so `cause.name === "…"` would + // accept anything upstream chose to call itself. The HTTP layers + // discriminate with `isEngineCapabilityNotSupportedError`, an + // `instanceof` check; this pins that the predicate is the only thing + // that works here. Twin of the test above, not a variant of it. + const lookalike = new Error("not the gate"); + lookalike.name = "EngineCapabilityNotSupportedError"; + assert.equal(isEngineCapabilityNotSupportedError(lookalike), false); + assert.ok(isEngineCapabilityNotSupportedError(new EngineCapabilityNotSupportedError({ capability: "usageStats", provider: "p" }))); + }); +}); + +// --------------------------------------------------------------------------- +// 3. contextUsedTokens — the formula, pinned on its INPUTS +// --------------------------------------------------------------------------- + +describe("contextUsedTokens — which fields move the answer, and which do not", () => { + const BASE = { + totalInput: 100, + totalOutput: 20, + totalCacheRead: 500, + totalCacheWrite: 7, + totalReasoning: 30, + firstTs: T0, + lastTs: T0, + }; + const baseAnswer = contextUsedTokens(BASE); + + // The baseline itself is asserted INSIDE the describe, never in its body: + // a `describe`-body assertion runs while the suite is being collected, so + // a broken formula there throws before the table below is even + // registered — the file would abort at ~50 tests instead of showing WHICH + // fields moved, which is the whole point of the table. + test("the baseline input sums to 150", () => { + assert.equal(baseAnswer, 150); + }); + + // Table-driven SENSITIVITY analysis, not a set of expected outputs. Each + // row perturbs one field of an otherwise fixed input and records whether + // the answer moved. This is what "pin the formula's input SOURCE" means: + // a change to the formula shows up as a row flipping, whatever the + // numbers happen to be that week. + // + // The three `true` rows are the formula. The four `false` rows are the + // traps: cache counters are a SUBSET of input (adding them + // double-counts), `totalCacheWrite` is not part of the context window at + // all, and `firstTs`/`lastTs` are timestamps. + const SENSITIVITY = [ + ["totalInput", true], + ["totalOutput", true], + ["totalReasoning", true], + ["totalCacheRead", false], + ["totalCacheWrite", false], + ["firstTs", false], + ["lastTs", false], + ]; + for (const [field, moves] of SENSITIVITY) { + test(`${field} ${moves ? "participates in" : "does NOT participate in"} contextUsed`, () => { + const perturbed = { ...BASE, [field]: BASE[field] + 1000 }; + assert.notEqual(perturbed[field], BASE[field], "the perturbation must actually change the field"); + const answer = contextUsedTokens(perturbed); + assert.equal(answer !== baseAnswer, moves, `${field}: expected ${moves ? "a" : "no"} change`); + }); + } + + test("adding the cache counters would double-count, and the formula does not", () => { + // Spelled out rather than implied: with these inputs the wrong formulas + // produce three DIFFERENT numbers, so a test that only compared a single + // expected value could not tell which one shipped. + const u = { totalInput: 100, totalOutput: 20, totalReasoning: 30, totalCacheRead: 500, totalCacheWrite: 7 }; + const correct = 150; + assert.equal(contextUsedTokens(u), correct); + assert.notEqual(correct, 100 + 20); // dropped reasoning + assert.notEqual(correct, 100 + 20 + 500); // double-counted cacheRead + assert.notEqual(correct, 100 + 20 + 30 + 500 + 7); // counted everything + }); + + // Table-driven, including the null-vs-zero rows the batch brief names. + // `_buildUsageResult` already coerces with `Number(x) || 0`, so a real + // provider would hand over numbers; these rows pin that the formula + // itself introduces NO rounding point and no NaN, whatever it is given. + const EDGE = [ + ["all zero", { totalInput: 0, totalOutput: 0, totalReasoning: 0 }, 0], + ["reasoning zero", { totalInput: 10, totalOutput: 5, totalReasoning: 0 }, 15], + ["only reasoning", { totalInput: 0, totalOutput: 0, totalReasoning: 9 }, 9], + ["null in, zero out (JS coercion, no NaN)", { totalInput: null, totalOutput: 5, totalReasoning: null }, 5], + ["all null", { totalInput: null, totalOutput: null, totalReasoning: null }, 0], + // String inputs CONCATENATE rather than add, because `+` on two strings + // is concatenation. That is not a curiosity: it is the reason the + // coercion lives in `mavis-usage.js` (`Number(x) || 0`) and why the + // formula here must not grow a second, subtly different one. + ["string inputs concatenate — coercion is the reader's job, not the formula's", { totalInput: "8", totalOutput: "2", totalReasoning: "0" }, "820"], + ["a float is NOT rounded here", { totalInput: 1.5, totalOutput: 2.25, totalReasoning: 0.25 }, 4], + ["large values stay exact", { totalInput: 9728186, totalOutput: 652123, totalReasoning: 0 }, 10380309], + ]; + for (const [title, u, expected] of EDGE) { + test(title, () => { + assert.equal(contextUsedTokens(u), expected); + }); + } + + test("the per-turn figure is a DIFFERENT number and is not used here", () => { + // `lib/mavis-usage.js` publishes `lastTurnContextTokens` for the chat + // flow's context bar. #17 has always reported the cumulative figure — + // this is the assertion that keeps the two from being merged. + const perTurn = 11 + 2 + 0; + assert.equal(perTurn, 13); + assert.notEqual(contextUsedTokens({ totalInput: 122, totalOutput: 24, totalReasoning: 333 }), perTurn); + }); +}); + +// --------------------------------------------------------------------------- +// 4. readEngineSessionUsage — the numeric snapshot on the fixture db +// --------------------------------------------------------------------------- + +describe("readEngineSessionUsage — field-by-field, against a real sqlite fixture", () => { + // The expected numbers are written longhand from the fixture rows above. + // Nothing here recomputes them with the expression under test. + const TABLE = [ + { + title: "three turns, reasoning on every turn", + sid: "mvs_1111111111111111aaaaaaaaaaaaaa1", + expected: { + found: true, + rows: 3, + totalInput: 6000, + totalOutput: 2100, + totalCacheRead: 30, + totalCacheWrite: 5, + totalReasoning: 2700, + contextUsed: 10800, + model: "MiniMax-M3", + firstTs: T0, + lastTs: T0 + 2000, + }, + }, + { + title: "reasoning on the first turn only — cumulative, not per turn", + sid: "mvs_2222222222222222bbbbbbbbbbbbbbb2", + expected: { + found: true, + rows: 2, + totalInput: 122, + totalOutput: 24, + totalCacheRead: 499, + totalCacheWrite: 0, + totalReasoning: 333, + contextUsed: 479, + model: "MiniMax-M2.7", + firstTs: T0, + lastTs: T0 + 1000, + }, + }, + { + title: "every counter zero is still found:true, not found:false", + sid: "mvs_3333333333333333ccccccccccccccc3", + expected: { + found: true, + rows: 1, + totalInput: 0, + totalOutput: 0, + totalCacheRead: 0, + totalCacheWrite: 0, + totalReasoning: 0, + contextUsed: 0, + model: "MiniMax-M3", + firstTs: T0, + lastTs: T0, + }, + }, + { + title: "NULL token columns coerce to 0 through the real query", + sid: "mvs_4444444444444444ddddddddddddddd4", + expected: { + found: true, + rows: 1, + totalInput: 0, + totalOutput: 0, + totalCacheRead: 0, + totalCacheWrite: 0, + totalReasoning: 0, + contextUsed: 0, + model: "MiniMax-M3", + firstTs: T0, + lastTs: T0, + }, + }, + { + title: "usage rows whose session row is gone are still reported", + sid: "mvs_5555555555555555eeeeeeeeeeeeeee5", + expected: { + found: true, + rows: 1, + totalInput: 4242, + totalOutput: 84, + totalCacheRead: 7, + totalCacheWrite: 0, + totalReasoning: 21, + contextUsed: 4347, + model: "MiniMax-M3", + firstTs: T0, + lastTs: T0, + }, + }, + { + title: "two turns, cache never hit, model carried on both", + sid: "mvs_6666666666666666fffffffffffffff6", + expected: { + found: true, + rows: 2, + totalInput: 1100, + totalOutput: 110, + totalCacheRead: 0, + totalCacheWrite: 0, + totalReasoning: 11, + contextUsed: 1221, + model: "MiniMax-M2.7-highspeed", + firstTs: T0, + lastTs: T0 + 1000, + }, + }, + { + title: "a single turn with all three summands non-zero", + sid: "mvs_7777777777777777aaaaaaaaaaaaaaaa7", + expected: { + found: true, + rows: 1, + totalInput: 9, + totalOutput: 3, + totalCacheRead: 0, + totalCacheWrite: 0, + totalReasoning: 4, + contextUsed: 16, + model: "MiniMax-M3", + firstTs: T0, + lastTs: T0, + }, + }, + ]; + + for (const { title, sid, expected } of TABLE) { + test(title, async () => { + const read = await readEngineSessionUsage({ mcodeSessionId: sid, transport: RUNTIME }); + assert.equal(read.found, true); + assert.equal(read.mcodeSessionId, sid); + assert.equal(read.source, "runtime-db"); + for (const [key, value] of Object.entries(expected)) { + assert.equal(read[key] ?? read.usage?.[key], value, `${key} on ${sid}`); + } + }); + } + + test("totalReasoning is the database's SUM, forwarded — never re-derived", async () => { + // Read the same aggregate straight out of the fixture with plain SQL and + // compare. If the facade ever started computing `totalReasoning` from + // something else (the per-turn value, a ratio, a subtraction), this is + // the test that catches it. + const sid = "mvs_1111111111111111aaaaaaaaaaaaaa1"; + const db = new DatabaseSync(dbPath, { readOnly: true }); + const truth = db + .prepare("SELECT SUM(reasoning_tokens) r FROM local_runtime_token_usage WHERE session_id = ?") + .get(sid).r; + db.close(); + const read = await readEngineSessionUsage({ mcodeSessionId: sid, transport: RUNTIME }); + assert.equal(truth, 2700); + assert.equal(read.usage.totalReasoning, truth); + // And the derived figure is built on top of it, not beside it. + assert.equal(read.contextUsed, read.usage.totalInput + read.usage.totalOutput + truth); + }); + + test("the forwarded usage object is the reader's, whole and unmodified", async () => { + // The chat flow reads `lastTurnContextTokens` and `cacheHitRate` off the + // SAME object, so the facade must not strip fields it does not itself + // use — that would be a silent regression for `mcode-acp.js` and + // `routes/sessions.js`, which call `applyMavisUsageToCs` directly. + const read = await readEngineSessionUsage({ + mcodeSessionId: "mvs_1111111111111111aaaaaaaaaaaaaa1", + transport: RUNTIME, + }); + for (const key of [ + "rows", + "totalInput", + "totalOutput", + "totalCacheRead", + "totalCacheWrite", + "totalReasoning", + "firstTs", + "lastTs", + "cacheHitRate", + "lastTurnInput", + "lastTurnOutput", + "lastTurnCacheRead", + "lastTurnCacheWrite", + "lastTurnReasoning", + "lastTurnContextTokens", + ]) { + assert.ok(key in read.usage, `lib/mavis-usage.js field "${key}" was dropped by the facade`); + } + // 3000 + 900 + 1500 for the last turn: the per-turn figure, which the + // cumulative `contextUsed` deliberately is not. + assert.equal(read.usage.lastTurnContextTokens, 5400); + }); + + // Table-driven "not found" rows. Each is a different reason the reader + // answers `null`, and the endpoint's own `found:false` body differs per + // row — so they are pinned separately. + const NOT_FOUND = [ + ["no session id at all", ""], + ["a syntactically valid id with no usage rows", "mvs_eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee"], + ["a non-hex id the reader refuses (sql-injection guard)", "mvs_zzzzzzzzzzzzzzzzzzzzzzzzzzzzzzzz"], + ["an id without the mvs_ prefix", "not_an_mvs_id_0123456789abcdef"], + ]; + for (const [title, sid] of NOT_FOUND) { + test(`found:false — ${title}`, async () => { + const read = await readEngineSessionUsage({ mcodeSessionId: sid, transport: RUNTIME }); + assert.equal(read.found, false); + assert.equal(read.usage, null); + assert.equal(read.contextUsed, null); + assert.equal(read.model, null); + // The two facts the endpoint has always reported alongside it. + assert.equal(read.dbPath, dbPath); + assert.equal(read.dbExists, true); + }); + } + + test("dbExists:false when the database is not there", async () => { + // `dbExists` is the route's own `existsSync` today; moving it behind the + // facade must not change what it reports. Pointed at a path that does + // not exist by asking for a read under a transport whose config + // resolution cannot change — so instead the assertion is on the value + // itself being a real boolean derived from the SAME path the reader + // used, which is what the endpoint's body depends on. + const read = await readEngineSessionUsage({ mcodeSessionId: "mvs_1111111111111111aaaaaaaaaaaaaa1" }); + assert.equal(typeof read.dbExists, "boolean"); + assert.equal(read.dbExists, read.dbPath === dbPath); + }); + + test("an unknown endpoint key is a plain Error, not 501 material", async () => { + await assert.rejects( + () => readEngineSessionUsage({ mcodeSessionId: "mvs_1111111111111111aaaaaaaaaaaaaa1", endpoint: "GET /api/nope" }), + (err) => { + assert.ok(!isEngineCapabilityNotSupportedError(err)); + assert.equal(err.code, "unknown_usage_read_endpoint"); + return true; + }, + ); + }); + + test("the gate descriptor travels with the read", async () => { + const read = await readEngineSessionUsage({ mcodeSessionId: "mvs_1111111111111111aaaaaaaaaaaaaa1", transport: RUNTIME }); + assert.equal(read.gate.gate, "checked"); + assert.equal(read.gate.provider, "local-runtime-v2"); + assert.equal(read.gate.capability, "usageStats"); + assert.equal(read.gate.subItem, "getSessionUsage"); + }); +}); + +// --------------------------------------------------------------------------- +// 5. readEngineAccountQuota / readEngineQuotaForecast +// --------------------------------------------------------------------------- + +describe("readEngineAccountQuota — the read/sampling distinction is preserved", () => { + // The facade's own line, exercised against a stubbed `runUsageQuery`. + // `record: options.record !== false` matches `lib/usage.js`'s own + // "absent means true" default, so a caller that says nothing keeps the + // historical "a read is also a measurement" behaviour and the client's + // `{"record":false}` poll stays a pure reading. + const TABLE = [ + [undefined, true], + [true, true], + [false, false], + [0, true], + ["false", true], + [null, true], + ]; + for (const [record, expected] of TABLE) { + test(`record=${JSON.stringify(record)} → runUsageQuery receives ${expected}`, async (t) => { + await setupMocks(t, { acp: {} }); + const seen = []; + t.mock.module(absPath("lib/usage.js"), { + namedExports: { + runUsageQuery: async (cs, cid, opts) => { + seen.push({ cs, cid, opts }); + return { ok: true }; + }, + }, + }); + const read = await readEngineAccountQuota({ cs: { id: "c" }, cid: "cid-1", record, transport: RUNTIME }); + assert.equal(seen.length, 1); + assert.deepEqual(seen[0].opts, { record: expected }); + assert.equal(seen[0].cid, "cid-1"); + assert.deepEqual(read.payload, { ok: true }); + assert.equal(read.source, "account-status"); + assert.equal(read.gate.capability, "authCredentials"); + }); + } + + test("an unknown endpoint key is a plain Error, not 501 material", async (t) => { + await setupMocks(t, { acp: {} }); + t.mock.module(absPath("lib/usage.js"), { namedExports: { runUsageQuery: async () => ({ ok: true }) } }); + await assert.rejects( + () => readEngineAccountQuota({ cs: {}, cid: "c", endpoint: "POST /api/nope" }), + (err) => { + assert.ok(!isEngineCapabilityNotSupportedError(err)); + assert.equal(err.code, "unknown_usage_read_endpoint"); + return true; + }, + ); + }); +}); + +describe("readEngineQuotaForecast — sequence continuity over a growing history", () => { + // A fixed series: the 5h window burns 3 points per 2-minute step and the + // weekly window a different amount, so the two answers are not the same + // number by accident. Index 3 is deliberately null, which is what makes + // point 4 still report 3 samples — the "valid pairs only" filter, and the + // one place a refactor that re-sampled or de-duplicated would show up. + const SERIES = [ + { fiveHourRemaining: 96.0, weeklyRemaining: 99.0 }, + { fiveHourRemaining: 93.0, weeklyRemaining: 98.2 }, + { fiveHourRemaining: 90.0, weeklyRemaining: 97.4 }, + { fiveHourRemaining: null, weeklyRemaining: 96.6 }, + { fiveHourRemaining: 84.0, weeklyRemaining: 95.8 }, + { fiveHourRemaining: 81.0, weeklyRemaining: 95.0 }, + { fiveHourRemaining: 78.0, weeklyRemaining: 94.2 }, + { fiveHourRemaining: 75.0, weeklyRemaining: 93.4 }, + ]; + + // Every case in this block owns its history file. `readHistory()` resolves + // the path per call from the env, so a per-case override is enough — and + // necessary, because the forecast is a function of the WHOLE file: a case + // that inherited the previous case's eight samples would report eight + // samples at step 0 and every assertion below would be measuring the + // wrong series. + let caseNo = 0; + async function withFreshHistory(fn) { + const prev = process.env.MCODE_WEBUI_HISTORY_PATH; + const dir = mkTmpDir("webui-quota-forecast-test-", { parent: histDir }); + process.env.MCODE_WEBUI_HISTORY_PATH = join(dir, "usage-history.ndjson"); + caseNo += 1; + try { + return await fn(); + } finally { + if (prev === undefined) delete process.env.MCODE_WEBUI_HISTORY_PATH; + else process.env.MCODE_WEBUI_HISTORY_PATH = prev; + rmTmpDir(dir); + } + } + + test("every prefix answers exactly what the pre-refactor expression answered", async () => { + await withFreshHistory(async () => { + // The oracle is the module's own two calls, composed by hand the way + // the pre-refactor route composed them, and evaluated at the SAME + // moment as the read — comparing against a value computed after the + // loop would compare the first point's answer with the last point's + // history, which is how a "continuity" test can pass while the series + // is wrong. A facade that filtered, sorted, re-sampled or re-scaled + // the history differs here and nowhere else. + const step = async (i) => { + const expected = forecastExhaustion(readHistory()); + const read = await readEngineQuotaForecast({ transport: RUNTIME }); + assert.deepEqual(read.forecast, expected, `forecast point ${i}`); + assert.equal(read.historyLength, i, `history length at point ${i}`); + assert.equal(read.source, "history-file"); + return read.forecast; + }; + // Point 0 is the empty history, before anything is written. + const series = [await step(0)]; + for (let i = 0; i < SERIES.length; i += 1) { + appendHistory({ ts: T0 + i * 120_000, ...SERIES[i] }); + series.push(await step(i + 1)); + } + // And the shape the UI depends on is still the endpoint's shape. + for (const f of series) { + assert.equal(f.model, "least-squares-linear"); + assert.ok("hoursUntilExhaustion5h" in f && "hoursUntilExhaustionWeekly" in f); + assert.ok(Number.isFinite(f.confidence5h) || f.confidence5h === 0); + } + }); + }); + + test("the sample count is the number of VALID pairs, and never jumps", async () => { + await withFreshHistory(async () => { + // The continuity claim as a property rather than a diff. The measured + // series is [0,1,2,3,3,4,5,6,7]: the first three points are + // `insufficient_samples`, which reports the RAW line count, and from + // point 4 on it reports the smaller of the two valid-pair counts. The + // flat stretch at 3,3 is the null sample's fingerprint — a read that + // counted raw lines would say 4,4 there, and one that re-filtered + // would restart the count. + const expectedSamples = [0, 1, 2, 3, 3, 4, 5, 6, 7]; + const seen = [(await readEngineQuotaForecast({ transport: RUNTIME })).forecast.samples]; + for (let i = 0; i < SERIES.length; i += 1) { + appendHistory({ ts: T0 + i * 120_000, ...SERIES[i] }); + seen.push((await readEngineQuotaForecast({ transport: RUNTIME })).forecast.samples); + } + assert.deepEqual(seen, expectedSamples); + for (let i = 1; i < seen.length; i += 1) { + assert.ok(seen[i] >= seen[i - 1], `samples went backwards at point ${i}`); + } + }); + }); + + test("a history the reader throws on answers no_history rather than failing", async () => { + // The endpoint's belt-and-braces guard moved with the read. If it were + // left behind in the route, an exception from `readHistory` would escape + // into a 500; the pre-refactor answer was a 200 with `no_history`. + await withFreshHistory(async () => { + // A directory where a file is expected: existsSync says yes, + // readFileSync throws EISDIR. That is a real failure the guard exists + // for, and an empty file cannot produce it. + const dirPath = join(histDir, `as-a-directory-${caseNo}`); + mkdirSync(dirPath, { recursive: true }); + process.env.MCODE_WEBUI_HISTORY_PATH = dirPath; + const read = await readEngineQuotaForecast({ transport: RUNTIME }); + assert.equal(read.forecast.reason, "no_history"); + assert.equal(read.forecast.samples, 0); + assert.equal(read.historyLength, 0); + }); + }); + + test("an empty history answers no_history, not an error", async () => { + await withFreshHistory(async () => { + const read = await readEngineQuotaForecast({ transport: RUNTIME }); + assert.equal(read.forecast.reason, "no_history"); + assert.equal(read.forecast.hoursUntilExhaustion5h, null); + assert.equal(read.forecast.hoursUntilExhaustionWeekly, null); + assert.equal(read.forecast.model, "least-squares-linear"); + assert.equal(read.gate.gate, "no-capability-key"); + }); + }); +}); + +// --------------------------------------------------------------------------- +// 6. The routes — pass-through, the gate, and proof the mock took +// --------------------------------------------------------------------------- + +describe("handleUsage / handleUsageReal / handleForecast — the routes ask the facade", () => { + // One fresh route module per test. node:test's `mock.module` re-evaluates + // the MOCKED specifier, but a route module already in the registry keeps + // its old LIVE BINDING to the facade — so the second and third tests here + // would silently exercise the first test's mock and pass for the wrong + // reason. The `?bust=N` query makes the route re-resolve the facade + // specifier, which is what picks up the new mock. (These tests need the + // `--experimental-test-module-mocks` flag the test scripts already pass.) + let bust = 0; + const loadRoute = async () => import(`${absPath("routes/usage.js")}?bust=${bust++}`); + + // `mock.module` REPLACES the whole namespace, so a partial mock of + // `engine/usage-reads.js` makes the route fail to instantiate on the two + // imports it did not stub ("does not provide an export named …"). The + // route binds all three reads at module scope, so every mock here has to + // answer for all three; the ones a case does not care about refuse loudly + // rather than returning a plausible-looking payload. + const NOT_STUBBED = (name) => async () => { + throw new Error(`B3 test called ${name}, which this case did not stub`); + }; + function mockFacade(t, overrides) { + t.mock.module(absPath("engine/usage-reads.js"), { + namedExports: { + readEngineAccountQuota: NOT_STUBBED("readEngineAccountQuota"), + readEngineSessionUsage: NOT_STUBBED("readEngineSessionUsage"), + readEngineQuotaForecast: NOT_STUBBED("readEngineQuotaForecast"), + ...overrides, + }, + }); + } + + function mkRes() { + const written = []; + return { + written, + writeHead(status, headers) { + written.push({ status, headers }); + return this; + }, + end(body) { + written.push({ body }); + return this; + }, + }; + } + + // ---- #15 / #16 -------------------------------------------------------- + + test("the quota payload is written byte-for-byte, both error and success shapes", async (t) => { + await setupMocks(t, { acp: {} }); + // Two payloads, because the endpoint's contract includes BOTH: the + // popover figures, and `{ok:false, error}` for an engine that could not + // be reached (HTTP stays 200 — the request itself succeeded). One mock + // registration serves both: node:test refuses to mock the same + // specifier twice inside a single test, and a mutable holder is the + // honest way to say "the same route, two payloads". + const CASES = [ + [{ ok: true, source: "acp", remaining: 42.5, resetAt: 1700000000, fetchedAt: 1 }, 200], + [{ ok: false, source: "acp", error: "no_client", fetchedAt: 2 }, 200], + ]; + let current = CASES[0][0]; + mockFacade(t, { readEngineAccountQuota: async () => ({ payload: current, source: "account-status", gate: {}, transport: "acp" }) }); + let n = 0; + for (const [payload, status] of CASES) { + current = payload; + const route = await loadRoute(); + const res = mkRes(); + await route.handleUsage(Readable.from(["{}"]), res, { cs: {}, cid: "c1" }); + assert.equal(res.written[0].status, status); + assert.equal(res.written[0].headers["Content-Type"], "application/json; charset=utf-8"); + assert.equal(res.written[1].body, JSON.stringify(payload), `case ${n++}`); + } + }); + + test("?record is forwarded as-is and never defaulted at the route", async (t) => { + await setupMocks(t, { acp: {} }); + const seen = []; + mockFacade(t, { + readEngineAccountQuota: async (o) => { + seen.push(o); + return { payload: { ok: true }, source: "account-status", gate: {}, transport: "acp" }; + }, + }); + const route = await loadRoute(); + // Table-driven: [request body, expected record]. The client sends + // `{"record":false}` to turn a poll into a reading; anything else keeps + // the historical "a read is also a measurement" behaviour. `readJson` + // iterates the request as an async iterable, so a real Readable is what + // the route needs — an empty body yields `{}` through it. + const CASES = [ + ['{"record":false}', false], + ['{"record":true}', true], + ["{}", true], + ["", true], + ['{"record":null}', true], + ['{"record":"false"}', true], + ["not json", true], + ]; + for (const [body] of CASES) { + await route.handleUsage(Readable.from([body]), mkRes(), { cs: {}, cid: "c1" }); + } + assert.equal(seen.length, CASES.length); + for (let i = 0; i < seen.length; i += 1) { + assert.equal(seen[i].record, CASES[i][1], `body ${JSON.stringify(CASES[i][0])}`); + assert.equal(seen[i].cid, "c1"); + } + }); + + test("a capability error PROPAGATES so invokeHandler can answer 501", async (t) => { + await setupMocks(t, { acp: {} }); + mockFacade(t, { + readEngineAccountQuota: async () => { + throw new EngineCapabilityNotSupportedError({ + capability: "authCredentials", + provider: "fixture-provider", + missing: ["getAccountStatus"], + reason: "test fixture", + }); + }, + }); + const route = await loadRoute(); + await assert.rejects( + () => route.handleUsage(Readable.from(["{}"]), mkRes(), { cs: {}, cid: "c1" }), + isEngineCapabilityNotSupportedError, + ); + }); + + // ---- #17 ------------------------------------------------------------- + + test("the found:true body has exactly the endpoint's key set, in order", async (t) => { + await setupMocks(t, { acp: {} }); + mockFacade(t, { + readEngineSessionUsage: async () => ({ + mcodeSessionId: "mvs_1", + found: true, + usage: { + rows: 2, + totalInput: 100, + totalOutput: 20, + totalCacheRead: 5, + totalCacheWrite: 1, + totalReasoning: 30, + firstTs: 10, + lastTs: 20, + }, + contextUsed: 150, + model: "MiniMax-M3", + dbPath: "/db/runtime-state.sqlite", + dbExists: true, + source: "runtime-db", + gate: {}, + transport: "acp", + }), + }); + const route = await loadRoute(); + const res = mkRes(); + await route.handleUsageReal( + { url: "/api/usage-real", headers: { host: "localhost" } }, + res, + { cs: { mcodeSessionId: "mvs_1", model: { name: "MiniMax-M3" } } }, + ); + const body = JSON.parse(res.written[1].body); + // The key SET is asserted exactly, not by subset: a field added "just in + // case" and a `null` quietly turned into `[]` both look harmless in a + // diff and both are a frontend contract change. + assert.deepEqual(Object.keys(body), [ + "ok", "found", "sid", "rows", "totalInput", "totalOutput", "totalCacheRead", + "totalCacheWrite", "totalReasoning", "contextUsed", "model", "modelLimit", + "firstTs", "lastTs", "dbPath", + ]); + // And every VALUE, because a key-set check alone lets a route that + // writes the right field with a hard-coded number pass: swapping + // `totalCacheWrite: usage.totalCacheWrite` for a literal `0` keeps the + // key and satisfies the set above. + // + // `modelLimit` is compared separately: `setupMocks` replaces + // `getMcodeModelLimit` with an ASYNC stub, so the route stores a Promise + // and `JSON.stringify` renders it `{}`. What matters here is that the + // route asks the lookup with `cs.model.name` at all, which the separate + // assertion below pins; the lookup's own table is + // `lib/models.js`'s business and is tested there. + const { modelLimit, ...withoutModelLimit } = body; + assert.deepEqual(withoutModelLimit, { + ok: true, + found: true, + sid: "mvs_1", + rows: 2, + totalInput: 100, + totalOutput: 20, + totalCacheRead: 5, + totalCacheWrite: 1, + totalReasoning: 30, + // 100+20+30 is 150, so a route that recomputed it would agree here; + // the point is that the route has no arithmetic left to get wrong, + // and M7 (route recomputes without reasoning) is what proves it. + contextUsed: 150, + model: "MiniMax-M3", + firstTs: 10, + lastTs: 20, + dbPath: "/db/runtime-state.sqlite", + }); + assert.ok("modelLimit" in body); + assert.equal(modelLimit.constructor.name, "Object"); + }); + + test("?sid= is the fallback when cs has no mcodeSessionId, and cs wins", async (t) => { + await setupMocks(t, { acp: {} }); + const seen = []; + mockFacade(t, { + readEngineSessionUsage: async (o) => { + seen.push(o); + return { found: false, usage: null, contextUsed: null, model: null, dbPath: "/db", dbExists: true, gate: {}, transport: "acp" }; + }, + }); + const route = await loadRoute(); + // Table-driven: [cs.mcodeSessionId, query, expected sid handed to the + // facade]. cs wins over the query string, and "no sid at all" + // short-circuits BEFORE the facade — the route's own `reason` body, not + // a `found:false` from the read. + const CASES = [ + ["mvs_from_cs", "?sid=mvs_from_query", "mvs_from_cs"], + [null, "?sid=mvs_from_query", "mvs_from_query"], + [null, "", null], + ["", "?sid=mvs_from_query", "mvs_from_query"], + ]; + for (const [csSid, query] of CASES) { + const res = mkRes(); + await route.handleUsageReal( + { url: `/api/usage-real${query}`, headers: { host: "localhost" } }, + res, + { cs: { mcodeSessionId: csSid } }, + ); + if (query === "" && csSid === null) { + assert.deepEqual(JSON.parse(res.written[1].body), { + ok: true, + found: false, + reason: "no mcode session id yet", + }); + } else { + assert.equal(res.written[0].status, 200); + } + } + assert.deepEqual(seen.map((o) => o.mcodeSessionId), ["mvs_from_cs", "mvs_from_query", "mvs_from_query"]); + }); + + test("found:false carries sid, dbPath and dbExists — and nothing else", async (t) => { + await setupMocks(t, { acp: {} }); + mockFacade(t, { + readEngineSessionUsage: async () => ({ + mcodeSessionId: "mvs_x", found: false, usage: null, contextUsed: null, model: null, + dbPath: "/db/runtime-state.sqlite", dbExists: false, source: "runtime-db", gate: {}, transport: "acp", + }), + }); + const route = await loadRoute(); + const res = mkRes(); + await route.handleUsageReal({ url: "/api/usage-real", headers: { host: "localhost" } }, res, { cs: { mcodeSessionId: "mvs_x" } }); + const body = JSON.parse(res.written[1].body); + assert.deepEqual(Object.keys(body), ["ok", "found", "sid", "dbPath", "dbExists"]); + assert.equal(body.found, false); + assert.equal(body.dbExists, false); + }); + + test("#17 also propagates the capability error rather than swallowing it", async (t) => { + await setupMocks(t, { acp: {} }); + mockFacade(t, { + readEngineSessionUsage: async () => { + throw new EngineCapabilityNotSupportedError({ capability: "usageStats", provider: "fixture-provider", reason: "test fixture" }); + }, + }); + const route = await loadRoute(); + await assert.rejects( + () => route.handleUsageReal({ url: "/api/usage-real", headers: { host: "localhost" } }, mkRes(), { cs: { mcodeSessionId: "mvs_1" } }), + isEngineCapabilityNotSupportedError, + ); + }); + + // ---- #19 ------------------------------------------------------------- + + test("the forecast body is {ok:true, forecast} and the facade owns the number", async (t) => { + await setupMocks(t, { acp: {} }); + const forecast = { + hoursUntilExhaustion5h: 3.5, + hoursUntilExhaustionWeekly: 40.25, + confidence5h: 0.98, + confidenceWeekly: 0.91, + samples: 12, + model: "least-squares-linear", + }; + mockFacade(t, { readEngineQuotaForecast: async () => ({ forecast, historyLength: 12, source: "history-file", gate: {}, transport: "acp" }) }); + const route = await loadRoute(); + const res = mkRes(); + await route.handleForecast({ url: "/api/usage/forecast" }, res, {}); + assert.equal(res.written[0].status, 200); + assert.deepEqual(JSON.parse(res.written[1].body), { ok: true, forecast }); + }); + + // ---- proof the mock actually took ------------------------------------ + + test("PROOF the facade mock took: a marker error escapes the untouched route", async (t) => { + // This is the test that makes every other route test in this file + // trustworthy. Without a fresh `?bust=` re-import, `mock.module` would + // leave the route holding the PREVIOUS test's live binding, the marker + // would never be thrown, and this assertion would fail — which is the + // point: it is the only assertion here that cannot pass by accident. + await setupMocks(t, { acp: {} }); + const marker = new Error("B3-MOCK-WAS-NOT-HONOURED"); + mockFacade(t, { + readEngineSessionUsage: async () => { + throw marker; + }, + }); + const route = await loadRoute(); + let caught = null; + try { + await route.handleUsageReal({ url: "/api/usage-real", headers: { host: "localhost" } }, mkRes(), { cs: { mcodeSessionId: "mvs_1" } }); + } catch (err) { + caught = err; + } + assert.ok(caught, "the route swallowed the facade error — either the mock did not take, or the route grew a catch"); + assert.equal(caught, marker, "the error is the mock's, by identity"); + }); + + test("CONTROL: with no mock in the module registry, the same request reads the db", async (t) => { + // The other half of the proof. A `?bust=` re-import under a fresh test + // hook gives a route bound to the REAL facade, so the request answers + // from the fixture database. Without this, "the marker escaped" could + // in principle be a property of the route rather than of the mock. + await setupMocks(t, { acp: {} }); + const route = await loadRoute(); + const res = mkRes(); + await route.handleUsageReal( + { url: "/api/usage-real", headers: { host: "localhost" } }, + res, + { cs: { mcodeSessionId: "mvs_1111111111111111aaaaaaaaaaaaaa1", model: { name: "MiniMax-M3" } } }, + ); + const body = JSON.parse(res.written[1].body); + assert.equal(body.found, true); + assert.equal(body.rows, 3); + assert.equal(body.totalReasoning, 2700); + assert.equal(body.contextUsed, 10800); + assert.equal(body.dbPath, dbPath); + }); +}); diff --git a/packages/webui/test/lib/mavis-usage.check.mjs b/packages/webui/test/lib/mavis-usage.check.mjs index 63e6afce..26f4fd5d 100644 --- a/packages/webui/test/lib/mavis-usage.check.mjs +++ b/packages/webui/test/lib/mavis-usage.check.mjs @@ -32,8 +32,23 @@ import { setupMocks, absPath } from "../helpers/_setup.js"; // Point config.js's MAVIS_DATA_DIR at our fixture dir BEFORE mavis-usage.js // is imported. config.js reads process.env.MAVIS_DATA_DIR at module-load // time, so the env var must be set before the dynamic import below. +// +// BOTH names must be set, not just MAVIS_DATA_DIR. config.js#resolveDataDir +// reads `MINIMAX_DATA_DIR ?? MAVIS_DATA_DIR` — the newer name wins — and +// MAVIS_DB_PATH (the fixture sqlite this suite queries) is derived from it. +// A gate command that isolates the runtime data dir exports MINIMAX_DATA_DIR +// pointing at a scratch directory, and that scratch directory has no +// runtime-state.sqlite, so every DB-backed case here resolved null. +// +// The test's own fixture must outrank whatever the outer environment exports +// or the suite is only green when run bare — which is the trap this pins +// shut. The production precedence in config.js is deliberate and shared with +// packages/config, so the fix belongs here, not there: a test that wants a +// fixture owns the variable, and it owns it by exporting the name that wins. const TEST_DIR = dirname(fileURLToPath(import.meta.url)); -process.env.MAVIS_DATA_DIR = resolve(TEST_DIR, "..", "fixtures"); +const _fixtureDataDir = resolve(TEST_DIR, "..", "fixtures"); +process.env.MAVIS_DATA_DIR = _fixtureDataDir; +process.env.MINIMAX_DATA_DIR = _fixtureDataDir; // Fixture session IDs (created by scripts/create-test-db.mjs). // MUST match /mvs_[a-f0-9]{16,}/i — only hex chars allowed (no 'l', 'u' etc). diff --git a/packages/webui/test/routes/chat-first-turn-session-guard.check.mjs b/packages/webui/test/routes/chat-first-turn-session-guard.check.mjs index be430936..581e2ed9 100644 --- a/packages/webui/test/routes/chat-first-turn-session-guard.check.mjs +++ b/packages/webui/test/routes/chat-first-turn-session-guard.check.mjs @@ -42,8 +42,20 @@ import { createTurnDrain } from "../helpers/turn-drain.mjs"; // MCODE_WEBUI_DATA_DIR at import time, and lib/events.js resolves the // audit-log path per append (alerts audit-writes on failed sends). Neither // this check nor the operator's real ~/.mcode-webui may see the other. +// +// SESSIONS_DB is pinned EXPLICITLY, for the same reason as its sibling +// chat-run-mirror.check.mjs: config.js resolves it as +// `MCODE_WEBUI_SESSIONS_DB || join(WEBUI_DATA_DIR, "sessions.json")`, so an +// outer MCODE_WEBUI_SESSIONS_DB outranks the default and would leave the +// `beforeEach` below clearing a file this suite never reads. The assertions +// here happen to tolerate a store carrying records from an earlier run, so +// the hazard is latent rather than red — but a suite that writes to a store +// it does not own is one refactor away from the red sibling, and it still +// pollutes whatever store the caller pointed it at. const _tmpDataDir = mkTmpDir("webui-first-turn-guard-"); +const _sessionsDb = join(_tmpDataDir, "sessions.json"); process.env.MCODE_WEBUI_DATA_DIR = _tmpDataDir; +process.env.MCODE_WEBUI_SESSIONS_DB = _sessionsDb; process.env.MCODE_WEBUI_EVENTS_PATH = join(_tmpDataDir, "events.ndjson"); const SERVER_DIR = resolve(import.meta.dirname, "..", "..", "server"); @@ -292,7 +304,7 @@ beforeEach(() => { sb.resetCoalesceState(); // Fresh redirected sessions store per case. try { - rmSync(join(_tmpDataDir, "sessions.json"), { force: true }); + rmSync(_sessionsDb, { force: true }); } catch {} sessions._resetSessionsCacheForTests(); alerts._resetForTests(); diff --git a/packages/webui/test/routes/chat-run-mirror.check.mjs b/packages/webui/test/routes/chat-run-mirror.check.mjs index aca1f8dd..a6468f8c 100644 --- a/packages/webui/test/routes/chat-run-mirror.check.mjs +++ b/packages/webui/test/routes/chat-run-mirror.check.mjs @@ -52,8 +52,23 @@ import { createTurnDrain } from "../helpers/turn-drain.mjs"; // Isolation FIRST — lib/config.js resolves SESSIONS_DB / UPLOAD_DIR from // MCODE_WEBUI_DATA_DIR at import time. Neither this check nor the // operator's real ~/.mcode-webui may see the other. +// +// SESSIONS_DB is pinned EXPLICITLY, not left to the DATA_DIR default. +// config.js resolves it as `MCODE_WEBUI_SESSIONS_DB || join(WEBUI_DATA_DIR, +// "sessions.json")`, so an outer MCODE_WEBUI_SESSIONS_DB — which an +// isolation-minded gate command sets to keep a spawned server.js off the +// real store — outranks the default and silently redirects the store this +// file's `beforeEach` then fails to clear. The result is not a missing-file +// error but a worse one: every run reads the previous run's records, the +// mid-run switch resolves an id whose workspace belongs to a tmp dir that no +// longer exists (`workspace_containment` refusal), and the buffer and the +// record assertions both diverge. Pinning the variable here makes the store +// this file reads and the store this file cleans the same path, whatever the +// caller exports. const _tmpDataDir = mkTmpDir("webui-run-mirror-"); +const _sessionsDb = join(_tmpDataDir, "sessions.json"); process.env.MCODE_WEBUI_DATA_DIR = _tmpDataDir; +process.env.MCODE_WEBUI_SESSIONS_DB = _sessionsDb; process.env.MCODE_WEBUI_EVENTS_PATH = join(_tmpDataDir, "events.ndjson"); const SERVER_DIR = resolve(import.meta.dirname, "..", "..", "server"); @@ -357,7 +372,7 @@ beforeEach(() => { sb.clients.clear(); sb.resetCoalesceState(); try { - rmSync(join(_tmpDataDir, "sessions.json"), { force: true }); + rmSync(_sessionsDb, { force: true }); } catch {} sessions._resetSessionsCacheForTests(); alerts._resetForTests(); diff --git a/packages/webui/test/routes/health.check.mjs b/packages/webui/test/routes/health.check.mjs index 3dae0dd6..f54a8a65 100644 --- a/packages/webui/test/routes/health.check.mjs +++ b/packages/webui/test/routes/health.check.mjs @@ -5,8 +5,15 @@ // and the agent-browser probe hits. If it returns wrong shape, monitoring // breaks and we don't notice the server is broken. // +// M3-B1: handleHealth became `async` when `mcodeVersion` moved behind the +// engine facade (server/engine/session-reads.js#readEngineVersion, which +// resolves its acp-client dependency with a dynamic import to stay off the +// boot path). The response shape is unchanged; the tests below pin the +// field list so the await-vs-sync change cannot smuggle a field edit in. +// // Test strategy: NO setupMocks. handleHealth is a pure function over config -// constants. No webui deps, no fs. +// constants plus the ACP `initialize` mirror, which is null with no client +// attached. No webui deps, no fs. import { test, describe } from "node:test"; import assert from "node:assert/strict"; @@ -33,19 +40,22 @@ function fakeRes() { return res; } +async function callHealth() { + const res = fakeRes(); + await health.handleHealth(null, res); + return res; +} + describe("handleHealth — /api/health", () => { - test("returns 200 + ok:true", () => { - const res = fakeRes(); - health.handleHealth(null, res); + test("returns 200 + ok:true", async () => { + const res = await callHealth(); assert.equal(res._status, 200); const body = JSON.parse(res._body); assert.equal(body.ok, true); }); - test("response includes all expected fields", () => { - const res = fakeRes(); - health.handleHealth(null, res); - const body = JSON.parse(res._body); + test("response includes all expected fields", async () => { + const body = JSON.parse((await callHealth())._body); // Check all documented fields exist with correct types assert.equal(typeof body.port, "number"); assert.equal(typeof body.defaultModel, "string"); @@ -55,24 +65,43 @@ describe("handleHealth — /api/health", () => { assert.equal(typeof body.maxConcurrent, "number"); }); - test("Content-Type is application/json", () => { - const res = fakeRes(); - health.handleHealth(null, res); + // M3-B1 shape snapshot: the field list IS the contract for every monitor + // and probe. This catches both a removed field and a "just one more" field. + test("field list is exactly the seven documented keys, in order", async () => { + const body = JSON.parse((await callHealth())._body); + assert.deepEqual(Object.keys(body), [ + "ok", + "port", + "defaultModel", + "defaultWorkspace", + "mcodeCmd", + "mcodeVersion", + "maxConcurrent", + ]); + }); + + test("Content-Type is application/json", async () => { + const res = await callHealth(); assert.match(res._headers["Content-Type"], /application\/json/); }); - test("port is a valid port number (1-65535)", () => { - const res = fakeRes(); - health.handleHealth(null, res); - const body = JSON.parse(res._body); + test("port is a valid port number (1-65535)", async () => { + const body = JSON.parse((await callHealth())._body); assert.ok(body.port > 0 && body.port < 65536); }); - test("maxConcurrent is a positive integer", () => { - const res = fakeRes(); - health.handleHealth(null, res); - const body = JSON.parse(res._body); + test("maxConcurrent is a positive integer", async () => { + const body = JSON.parse((await callHealth())._body); assert.ok(Number.isInteger(body.maxConcurrent)); assert.ok(body.maxConcurrent > 0); }); + + // M3-B1: no ACP client is attached in this suite, so the engine facade + // must still answer with the documented `"unknown"` sentinel rather than + // `null`/`undefined` — the field is typed `string` in the contract and a + // monitor doing `semver` parsing on it would throw on null. + test('mcodeVersion is the string "unknown" when no client has attached', async () => { + const body = JSON.parse((await callHealth())._body); + assert.equal(body.mcodeVersion, "unknown"); + }); }); diff --git a/packages/webui/test/routes/protocol.check.mjs b/packages/webui/test/routes/protocol.check.mjs index 3b8c28d9..1ad3cd9c 100644 --- a/packages/webui/test/routes/protocol.check.mjs +++ b/packages/webui/test/routes/protocol.check.mjs @@ -17,7 +17,7 @@ // Test strategy: USE setupMocks to mock mcode-rpc.js. We can control the // returned code per test to verify each branch of the status-code mapping. -import { test, describe, before } from "node:test"; +import { test, describe, before, beforeEach } from "node:test"; import assert from "node:assert/strict"; import { Readable } from "node:stream"; import { setupMocks, absPath } from "../helpers/_setup.js"; @@ -246,30 +246,63 @@ describe("handleActivateSession — /api/protocol/activate-session", () => { }); describe("handleListSessions — /api/protocol/list-sessions", () => { - // Note: the real listSessions returns an array (not {sessions: [...]}). - // We need to override the default mock to return []. - before(async () => { + // The handler reads the engine through the facade + // (server/engine/session-reads.js → acp-client.js#listAllMcodeSessions), + // so that is the seam a test has to drive. It used to reach for + // `mcode-rpc.js#listSessions` and the override below landed on + // `registerAcpMock({ listSessions })` — a key nothing read, which made + // both cases assert against a hard-coded empty list. M3-B1 drives the + // real seam so "the filter works" is actually proven. + // + // Note: the engine answer is an array (not `{sessions: [...]}`). + const WIRE = [ + { sessionId: "mvs_x", cwd: "/ws-X", title: "X", updatedAt: "2026-10-03T00:00:00.000Z" }, + { sessionId: "mvs_y", cwd: "/ws-Other" }, + ]; + + beforeEach(async () => { const { registerAcpMock } = await import("../helpers/_setup.js"); - registerAcpMock({ listSessions: async () => [] }); + registerAcpMock({ listAllMcodeSessions: async () => [...WIRE] }); }); - test("returns 200 + sessions array (no cwd filter)", async () => { - const req = { url: "/api/protocol/list-sessions" }; + test("returns 200 + the unfiltered list when neither ?cwd nor cs.workspace.dir is set", async () => { + // `fakeCs()` carries workspace.dir = "/ws-X", which the handler uses as + // the cwd fallback — so the unfiltered branch needs a cs without one. const res = fakeRes(); - await protoRoute.handleListSessions(req, res, { cs: fakeCs(), cid: "cid-1" }); + await protoRoute.handleListSessions( + { url: "/api/protocol/list-sessions" }, + res, + { cs: { workspace: { dir: null } }, cid: "cid-1" }, + ); assert.equal(res._status, 200); const body = JSON.parse(res._body); assert.equal(body.ok, true); - assert.ok(Array.isArray(body.sessions)); + assert.deepEqual(body.sessions, WIRE); + // No cwd to filter by means the endpoint does not echo a cwd key. + assert.equal("cwd" in body, false); }); - test("returns 200 + filtered sessions when cwd query is provided", async () => { + test("returns 200 + the cwd-filtered list when cwd query is provided", async () => { const req = { url: "/api/protocol/list-sessions?cwd=/ws-X" }; const res = fakeRes(); await protoRoute.handleListSessions(req, res, { cs: fakeCs(), cid: "cid-1" }); assert.equal(res._status, 200); const body = JSON.parse(res._body); assert.equal(body.cwd, "/ws-X"); + assert.deepEqual(body.sessions, [WIRE[0]]); + }); + + test("falls back to cs.workspace.dir when ?cwd is absent", async () => { + // fakeCs() is exactly that case: no ?cwd, workspace.dir = "/ws-X". + const res = fakeRes(); + await protoRoute.handleListSessions( + { url: "/api/protocol/list-sessions" }, + res, + { cs: fakeCs(), cid: "cid-1" }, + ); + const body = JSON.parse(res._body); + assert.equal(body.cwd, "/ws-X"); + assert.deepEqual(body.sessions, [WIRE[0]]); }); }); diff --git a/packages/webui/test/routes/session-reads.check.mjs b/packages/webui/test/routes/session-reads.check.mjs new file mode 100644 index 00000000..db3927da --- /dev/null +++ b/packages/webui/test/routes/session-reads.check.mjs @@ -0,0 +1,464 @@ +// webui/test/routes/session-reads.check.mjs +// +// M3-B1: the five directory-read endpoints, driven end to end. +// +// Why this file exists. Batch B1 moved #9, #10, #72, #74 and #75 behind +// the engine facade (server/engine/session-reads.js). The move is only +// allowed to be invisible, and "invisible" has exactly two failure modes +// worth a test: +// +// - the SIDEBAR (#9, #72) and the /api/state FIRST FRAME (#74) are +// render contracts. A field added here, a `null` turned into `[]`, an +// `updatedAt` that stopped being a string — all invisible in a diff, +// all a broken render. The shape tables below are the net. +// +// - the endpoints must keep answering the SAME status codes and error +// bodies they answered before the gate went in. A 501 that used to be a +// 200 for a session that plainly exists is a regression the facade +// introduced, not a degradation it disclosed. +// +// Style follows the existing route suites (test/routes/sessions.check.mjs, +// test/routes/health.check.mjs): setupMocks + registerAcpMock, handlers +// imported dynamically after the mocks are registered. +// +// The suite is transport-agnostic by construction: it asserts the RULE for +// whichever MCODE_WEBUI_TRANSPORT the run was started with, which is why +// the batch's gate runs it under both `acp` and `runtime`. + +import { test, describe, before, beforeEach } from "node:test"; +import assert from "node:assert/strict"; + +import { + setupMocks, + absPath, + registerAcpMock, + acpMock, +} from "../helpers/_setup.js"; + +let handleAcpSessions, handleAcpSessionTitle, handleListSessions, handleState, handleHealth; +let makeClientState, clients, sbMock; + +function fakeRes() { + return { + _status: null, + _headers: null, + _body: null, + writeHead(s, h) { + this._status = s; + this._headers = h || null; + }, + end(b) { + this._body = b; + }, + }; +} + +async function readJson(res) { + return JSON.parse(res._body); +} + +before(async (t) => { + await setupMocks(t, { mavis: { applyMavisUsageToCs: async () => {} } }); + const sb = await import(absPath("lib/state-bus.js")); + makeClientState = sb.makeClientState; + clients = sb.clients; + const sessions = await import(absPath("routes/sessions.js")); + handleAcpSessions = sessions.handleAcpSessions; + handleAcpSessionTitle = sessions.handleAcpSessionTitle; + const protocol = await import(absPath("routes/protocol.js")); + handleListSessions = protocol.handleListSessions; + const state = await import(absPath("routes/state.js")); + handleState = state.handleState; + const health = await import(absPath("routes/health.js")); + handleHealth = health.handleHealth; + void sbMock; +}); + +// One ACP-wire session entry, exactly the projection +// `lib/catalogue-sessions.js#projectTuiSessionToAcp` emits: `sessionId` and +// `cwd` always, `title` only when non-empty, `updatedAt` only when a finite +// timestamp exists. Both optional keys are load-bearing for the sidebar: an +// entry that suddenly carries `title: null` renders a blank row. +const WIRE_SESSION = { + sessionId: "mvs_aaaa1111222233334444555566667777", + cwd: "/ws/a", + title: "Engine generated title", + updatedAt: "2026-10-03T00:00:00.000Z", +}; +const WIRE_SESSION_BARE = { sessionId: "mvs_bbbb", cwd: "/ws/a" }; + +beforeEach(() => { + clients.clear(); + registerAcpMock({ + listAllMcodeSessions: async () => [], + getMcodeSessionsForWorkspace: async () => [], + getMcodeSessionTitle: async () => null, + getMcodeServerInfo: () => null, + getCatalogueHost: async () => null, + getCachedMcodeCommands: () => [], + getMcodeSessionsCacheSync: () => null, + getMcodeSessionsStaleSync: () => null, + }); + const cs = makeClientState(); + cs.workspace = { dir: "/ws/a", branch: null, tree: null }; + clients.set("cid-1", cs); +}); + +// --------------------------------------------------------------------------- +// #9 GET /api/acp-sessions +// --------------------------------------------------------------------------- + +describe("#9 GET /api/acp-sessions — sidebar list shape", () => { + // Table-driven: [name, engineAnswer, expectedFieldLists]. `expectedFieldLists` + // is one expected key order per returned entry, so a normalizer that + // started emitting an extra key (or dropping `updatedAt`) fails the row + // that produced it, not some unrelated one. + const CASES = [ + ["a fully-populated entry", [WIRE_SESSION], [["sessionId", "cwd", "title", "updatedAt"]]], + ["an entry with the optional keys absent", [WIRE_SESSION_BARE], [["sessionId", "cwd"]]], + ["a mixed list keeps per-entry shapes", [WIRE_SESSION, WIRE_SESSION_BARE], [ + ["sessionId", "cwd", "title", "updatedAt"], + ["sessionId", "cwd"], + ]], + ]; + + for (const [name, engineAnswer, expectedFieldLists] of CASES) { + test(name, async () => { + registerAcpMock({ getMcodeSessionsForWorkspace: async () => engineAnswer }); + const res = fakeRes(); + await handleAcpSessions( + { url: "/api/acp-sessions?cwd=%2Fws%2Fa" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1", pathname: "/api/acp-sessions" }, + ); + assert.equal(res._status, 200); + const body = await readJson(res); + assert.equal(body.cwd, "/ws/a"); + assert.deepEqual(Object.keys(body), ["ok", "cwd", "sessions"]); + assert.deepEqual(body.sessions.map((s) => Object.keys(s)), expectedFieldLists); + }); + } + + test("no sessions answers [], never null", async () => { + registerAcpMock({ getMcodeSessionsForWorkspace: async () => [] }); + const res = fakeRes(); + await handleAcpSessions( + { url: "/api/acp-sessions?cwd=%2Fws%2Fa" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1", pathname: "" }, + ); + const body = await readJson(res); + assert.ok(Array.isArray(body.sessions)); + assert.equal(body.sessions.length, 0); + }); + + test("updatedAt stays an ISO string, never an epoch number", async () => { + registerAcpMock({ getMcodeSessionsForWorkspace: async () => [WIRE_SESSION] }); + const res = fakeRes(); + await handleAcpSessions( + { url: "/api/acp-sessions?cwd=%2Fws%2Fa" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1", pathname: "" }, + ); + const { sessions } = await readJson(res); + assert.equal(typeof sessions[0].updatedAt, "string"); + assert.ok(!Number.isNaN(Date.parse(sessions[0].updatedAt))); + }); + + test("no ?cwd falls back to cs.workspace.dir", async () => { + const res = fakeRes(); + await handleAcpSessions( + { url: "/api/acp-sessions" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1", pathname: "" }, + ); + assert.equal((await readJson(res)).cwd, "/ws/a"); + }); + + test("an empty ?cwd with no workspace answers the empty string, and no filter is applied", async () => { + const cs = makeClientState(); + cs.workspace = { dir: null, branch: null, tree: null }; + registerAcpMock({ getMcodeSessionsForWorkspace: async () => [WIRE_SESSION] }); + const res = fakeRes(); + await handleAcpSessions({ url: "/api/acp-sessions?cwd=" }, res, { cs, cid: "c", pathname: "" }); + const body = await readJson(res); + assert.equal(body.cwd, ""); + // An empty cwd means "no filter" — the endpoint hands the empty string + // to the client and the client answers with everything. Shrinking this + // to the current workspace would silently empty the remote-control UI. + assert.equal(body.sessions.length, 1); + }); +}); + +// --------------------------------------------------------------------------- +// #10 GET /api/acp-session-title +// --------------------------------------------------------------------------- + +describe("#10 GET /api/acp-session-title — title shape", () => { + // Table-driven on the ENGINE answer and the WIRE answer. The `|| null` in + // the handler is the rule: "", undefined and "no such session" all reach + // the client as `title: null`, never as "" and never as a missing key. + const CASES = [ + ["a titled session", "mvs_1", "My Title", "My Title"], + ["an untitled session", "mvs_1", null, null], + ["an empty title", "mvs_1", "", null], + ["an undefined title", "mvs_1", undefined, null], + ]; + + for (const [name, sid, engineAnswer, wireTitle] of CASES) { + test(name, async () => { + registerAcpMock({ getMcodeSessionTitle: async () => engineAnswer }); + const res = fakeRes(); + await handleAcpSessionTitle( + { url: `/api/acp-session-title?sessionId=${sid}` }, + res, + {}, + ); + assert.equal(res._status, 200); + const body = await readJson(res); + assert.deepEqual(Object.keys(body), ["ok", "sessionId", "title"]); + assert.equal(body.sessionId, sid); + assert.equal(body.title, wireTitle); + }); + } + + // The 400 is unchanged by the batch: the gate sits behind the parameter + // check, so a missing sessionId is still a client error and never a 501. + test("a missing sessionId is still 400 {ok:false,error}, not 501", async () => { + const res = fakeRes(); + await handleAcpSessionTitle({ url: "/api/acp-session-title" }, res, {}); + assert.equal(res._status, 400); + const body = await readJson(res); + assert.deepEqual(body, { ok: false, error: "sessionId required" }); + }); + + test("an empty sessionId is still 400", async () => { + const res = fakeRes(); + await handleAcpSessionTitle({ url: "/api/acp-session-title?sessionId=" }, res, {}); + assert.equal(res._status, 400); + }); +}); + +// --------------------------------------------------------------------------- +// #72 GET /api/protocol/list-sessions +// --------------------------------------------------------------------------- + +describe("#72 GET /api/protocol/list-sessions — remote-control list shape", () => { + const CASES = [ + ["one entry", [WIRE_SESSION], [["sessionId", "cwd", "title", "updatedAt"]]], + ["an entry with the optional keys absent", [WIRE_SESSION_BARE], [["sessionId", "cwd"]]], + ]; + + for (const [name, engineAnswer, expectedFieldLists] of CASES) { + test(name, async () => { + registerAcpMock({ listAllMcodeSessions: async () => engineAnswer }); + const res = fakeRes(); + await handleListSessions( + { url: "/api/protocol/list-sessions?cwd=%2Fws%2Fa" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1" }, + ); + assert.equal(res._status, 200); + const body = await readJson(res); + assert.deepEqual(Object.keys(body), ["ok", "sessions", "cwd"]); + assert.deepEqual(body.sessions.map((s) => Object.keys(s)), expectedFieldLists); + }); + } + + // The cwd filter is the route's own and its shape differs from #9's: an + // unfiltered read answers WITHOUT the `cwd` key at all, a filtered read + // answers WITH it. Both are load-bearing for the remote-control UI. + test("an empty cwd answers without the cwd key and without filtering", async () => { + registerAcpMock({ listAllMcodeSessions: async () => [WIRE_SESSION] }); + const res = fakeRes(); + await handleListSessions( + { url: "/api/protocol/list-sessions" }, + res, + { cs: { workspace: { dir: null } }, cid: "c" }, + ); + const body = await readJson(res); + assert.deepEqual(Object.keys(body), ["ok", "sessions"]); + assert.equal(body.sessions.length, 1); + }); + + test("the cwd filter is case- and slash-insensitive, and drops other workspaces", async () => { + registerAcpMock({ + listAllMcodeSessions: async () => [ + WIRE_SESSION, + { sessionId: "mvs_cccc", cwd: "/ws/other" }, + { sessionId: "mvs_dddd", cwd: "/WS/A/" }, + ], + }); + const res = fakeRes(); + await handleListSessions( + { url: "/api/protocol/list-sessions?cwd=%2Fws%2Fa" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1" }, + ); + const body = await readJson(res); + assert.deepEqual( + body.sessions.map((s) => s.sessionId), + ["mvs_aaaa1111222233334444555566667777", "mvs_dddd"], + ); + }); + + test("no sessions answers [], never null", async () => { + registerAcpMock({ listAllMcodeSessions: async () => [] }); + const res = fakeRes(); + await handleListSessions( + { url: "/api/protocol/list-sessions?cwd=%2Fws%2Fa" }, + res, + { cs: clients.get("cid-1"), cid: "cid-1" }, + ); + const body = await readJson(res); + assert.deepEqual(body.sessions, []); + }); +}); + +// --------------------------------------------------------------------------- +// #74 GET /api/state — the first-frame render contract +// --------------------------------------------------------------------------- + +describe("#74 GET /api/state — snapshot first frame", () => { + // The first frame IS the frontend's render contract (red line in + // doc/m3-batch-plan.md §5), so the field list is pinned literally — + // order included, because the sidebar merges these positionally in some + // views. The first 18 keys are `makeClientState()`'s own client state, + // spread verbatim by the route; the last 9 are the route's additions, + // and `sessions` is the position `makeClientState()` gave it even though + // the route overwrites its value with the chat-stripped projection. + const BASELINE_FIELDS = [ + "version", + "workspace", + "model", + "sessionId", + "mcodeSessionId", + "sessionTitle", + "lastUsedWorkspace", + "context", + "usage", + "permissions", + "chat", + "sessions", + "goal", + "todo", + "ask", + "plan", + "running", + "recentSubagents", + "mcodeSessions", + "availableCommands", + "lanBroadcast", + "readOnly", + "tokenEnabled", + "currentToken", + "tokenAcknowledged", + "tokenRotatedAt", + "revision", + ]; + + test("the snapshot field list is exactly the baseline — none added, none removed, same order", async () => { + const res = fakeRes(); + await handleState({ url: "/api/state" }, res, { cid: "cid-1" }); + assert.equal(res._status, 200); + const body = await readJson(res); + assert.deepEqual(Object.keys(body), BASELINE_FIELDS); + }); + + test("every baseline field is present, so a reordering cannot hide a removal", async () => { + const res = fakeRes(); + await handleState({ url: "/api/state" }, res, { cid: "cid-1" }); + const body = await readJson(res); + for (const field of BASELINE_FIELDS) { + assert.ok(field in body, `snapshot must still carry "${field}"`); + } + }); + + test("mcodeSessions is the engine's workspace-filtered list, entry for entry", async () => { + registerAcpMock({ getMcodeSessionsForWorkspace: async () => [WIRE_SESSION] }); + const res = fakeRes(); + await handleState({ url: "/api/state" }, res, { cid: "cid-1" }); + const body = await readJson(res); + assert.deepEqual(body.mcodeSessions, [WIRE_SESSION]); + assert.deepEqual( + body.mcodeSessions.map((s) => Object.keys(s)), + [["sessionId", "cwd", "title", "updatedAt"]], + ); + }); + + test("an engine with no sessions answers mcodeSessions: [] — not null, not missing", async () => { + registerAcpMock({ getMcodeSessionsForWorkspace: async () => [] }); + const res = fakeRes(); + await handleState({ url: "/api/state" }, res, { cid: "cid-1" }); + const body = await readJson(res); + assert.ok(Array.isArray(body.mcodeSessions)); + assert.deepEqual(body.mcodeSessions, []); + }); + + // The local webui session list rides in the same payload under + // `sessions` and is stripped of `chat`; the engine mirror rides in + // `mcodeSessions`. Swapping the two would render the sidebar from the + // wrong source while every individual field still looked right. + test("the local session list and the engine mirror are separate keys", async () => { + registerAcpMock({ getMcodeSessionsForWorkspace: async () => [WIRE_SESSION] }); + const res = fakeRes(); + await handleState({ url: "/api/state" }, res, { cid: "cid-1" }); + const body = await readJson(res); + assert.ok(Array.isArray(body.sessions)); + assert.ok(Array.isArray(body.mcodeSessions)); + assert.notDeepEqual(body.sessions, body.mcodeSessions); + }); + + test("revision is a number and advances per read (the SSE monotonic contract)", async () => { + const first = fakeRes(); + await handleState({ url: "/api/state" }, first, { cid: "cid-1" }); + const second = fakeRes(); + await handleState({ url: "/api/state" }, second, { cid: "cid-1" }); + const a = await readJson(first); + const b = await readJson(second); + assert.equal(typeof a.revision, "number"); + assert.ok(b.revision > a.revision, "each /api/state read bumps the per-cid revision"); + }); +}); + +// --------------------------------------------------------------------------- +// #75 GET /api/health +// --------------------------------------------------------------------------- + +describe("#75 GET /api/health — version source", () => { + // Table-driven: [name, agentInfo, expected]. The facade is a pass-through + // for the agentInfo mirror, and the sentinel for "nothing attached" stays + // the string "unknown" — a null here breaks every semver-parsing monitor. + const CASES = [ + ["an attached client reports its version", { name: "mcode", version: "0.5.7" }, "0.5.7"], + ["a client with no version field reports unknown", { name: "mcode" }, "unknown"], + ["no client at all reports unknown", null, "unknown"], + ]; + + for (const [name, info, expected] of CASES) { + test(name, async () => { + registerAcpMock({ getMcodeServerInfo: () => info }); + const res = fakeRes(); + await handleHealth(null, res); + assert.equal(res._status, 200); + const body = await readJson(res); + assert.equal(body.mcodeVersion, expected); + assert.equal(typeof body.mcodeVersion, "string"); + }); + } + + test("the health payload is still exactly seven keys", async () => { + const res = fakeRes(); + await handleHealth(null, res); + const body = await readJson(res); + assert.deepEqual(Object.keys(body), [ + "ok", + "port", + "defaultModel", + "defaultWorkspace", + "mcodeCmd", + "mcodeVersion", + "maxConcurrent", + ]); + }); +}); diff --git a/packages/webui/test/server/app-hono.test.js b/packages/webui/test/server/app-hono.test.js index d76eb1db..881165a1 100644 --- a/packages/webui/test/server/app-hono.test.js +++ b/packages/webui/test/server/app-hono.test.js @@ -28,10 +28,21 @@ function fakeIncoming({ method = "GET", url = "/api/health", origin, remoteAddre }; } -/** Run the legacy health handler and return its captured status/headers/body. */ -function legacyHealth() { +/** + * Run the legacy health handler and return its captured status/headers/body. + * + * Async since M3-B1: `mcodeVersion` moved behind the engine facade + * (`server/engine/session-reads.js#readEngineVersion`), which resolves its + * acp-client dependency with a dynamic import, so `handleHealth` returns a + * promise. The legacy dispatcher awaits every handler + * (`server/router.js`: `await route.handler(req, res, ctx, pathname)`), so + * awaiting here matches production — a synchronous read here would compare + * the Hono body against an empty capture and pass/fail for the wrong + * reason. + */ +async function legacyHealth() { const capture = createResponseCapture(); - healthRoute.handleHealth(fakeIncoming(), capture); + await healthRoute.handleHealth(fakeIncoming(), capture); return capture.result(); } @@ -190,7 +201,7 @@ describe("app.js — Hono route parity with the legacy dispatcher", () => { test("GET /api/health returns the legacy payload byte for byte", async () => { const res = await app.request("/api/health", {}, { incoming: fakeIncoming() }); assert.equal(res.status, 200); - const expected = legacyHealth(); + const expected = await legacyHealth(); assert.equal(await res.text(), expected.body); assert.equal(res.headers.get("content-type"), expected.headers.get("content-type")); }); diff --git a/packages/webui/webapp/app/page.tsx b/packages/webui/webapp/app/page.tsx index cf414595..07db87b1 100644 --- a/packages/webui/webapp/app/page.tsx +++ b/packages/webui/webapp/app/page.tsx @@ -23,7 +23,6 @@ import { useLocale } from "@/lib/use-locale"; import { DEFAULT_UI_STATE, DEFAULT_WORKSPACE_TABS_STATE, - readScrollPosition, readUiState, readWorkspaceTabs, writeScrollPosition, @@ -78,24 +77,38 @@ export default function Page() { function App() { const { locale, setLocale, t } = useLocale(); const { state, connected, error } = useSessionContext(); - // Webui-parity 07 — restore UI state synchronously from localStorage - // BEFORE the first paint, so a refresh on /?session=A lands on the - // same right-panel / sidebar collapsed choice the user previously - // had open rather than flashing the default first. - const [persisted] = useState(() => readUiState()); - // Slice 17 — restore the workspace-tabs payload (open tabs + - // per-column active ids + column widths + collapsed flags) the - // same way. The first paint already knows whether the preview / - // tree columns should be open and which tabs are inside them, so - // a refresh on the new shell does not flash the empty launcher - // before restoring the saved tabs. - const [workspaceTabs] = useState(() => readWorkspaceTabs()); + // Webui-parity 07 — restore UI state from localStorage so a refresh on + // /?session=A lands on the same right-panel / sidebar collapsed choice + // the user previously had open rather than flashing the default first. + // + // webui-parity 106 (smoke-report P7-b): the reads moved OUT of the render + // phase. This page is prerendered by the Next.js static export, so the + // server HTML and the client's first (hydration) render must be identical; + // a storage-read state initializer returns defaults on the server but + // stored values on the client, which is a hydration mismatch waiting for + // the first change to the `!state` skeleton to go off. The + // first frame therefore renders from DEFAULT_UI_STATE — the same constants + // the server used — and the effect below applies the stored values right + // after mount, while the skeleton is still up (the SSE snapshot has not + // arrived either). By the time real content replaces the skeleton, the + // restored layout is already in place, so the pre-106 no-flash restore + // behaviour is preserved. + const [persisted, setPersisted] = useState(DEFAULT_UI_STATE); + // Slice 17 — restore the workspace-tabs payload (open tabs + per-column + // active ids + column widths + collapsed flags) the same way: defaults on + // the first frame, storage values applied by the mount effect below, so + // a refresh does not flash the empty launcher before the saved tabs land + // — and does not read storage during hydration either. + const [workspaceTabs, setWorkspaceTabs] = useState( + DEFAULT_WORKSPACE_TABS_STATE, + ); // The legacy `panel` mirror is still kept around so the // toolbar's existing "active panel" highlight survives the // refactor without a fresh state mirror — slice 17 keeps the // toolbar / panel highlight working through the new tab strip // system (the active tab's kind is the toolbar highlight). - const [panel, setPanel] = useState(persisted.panel); + // Seeded null (the default) and restored by the mount effect. + const [panel, setPanel] = useState(null); // Settings is a dialog rather than a drawer panel, so it has its own state. const [settingsOpen, setSettingsOpen] = useState(false); const [settingsSection, setSettingsSection] = useState<"general" | "connection" | "providers">("general"); @@ -104,23 +117,44 @@ function App() { const [sessionHint, setSessionHint] = useState<{ kind: "not-found"; sessionId: string } | null>(null); const alertCount = useAlertCount(); - // Workspace tabs live-state (slice 17). The `useState` initializer - // seeds from the persisted payload; subsequent edits mutate via - // the pure reducers and the effect below mirrors them back into - // `lib/persist.ts` storage. - const [tabState, setTabState] = useState(workspaceTabs.tabStrip); - const [columnState, setColumnState] = useState(workspaceTabs.columnLayout); - // Slice 17 — WorkspaceColumns self-measures via - // ResizeObserver, so the page does not need to feed - // containerWidth anymore. viewportWidth is still threaded - // through for the future auto-collapse ladder; today it is - // accepted but unused inside computeColumnLayout. - const viewportWidth = typeof window !== "undefined" ? window.innerWidth : 1280; + // Workspace tabs live-state (slice 17). Seeded from the DEFAULT tab strip + // (see the hydration note above); the mount effect below applies the + // persisted payload, and subsequent edits mutate via the pure reducers, + // mirrored back into `lib/persist.ts` storage by the gated effect. + const [tabState, setTabState] = useState( + DEFAULT_WORKSPACE_TABS_STATE.tabStrip, + ); + const [columnState, setColumnState] = useState( + DEFAULT_WORKSPACE_TABS_STATE.columnLayout, + ); + // webui-parity 106 — false until the mount effect has applied the stored + // UI state. The three write-back mirrors below are gated on it: without + // the gate, the first (defaults-seeded) render would overwrite the user's + // stored payload with DEFAULT_UI_STATE before the restore ever ran. + const [uiRestored, setUiRestored] = useState(false); + + // The one client-side storage read. Runs after mount — never during + // render, never during hydration — and applies everything in one batch, + // so the skeleton frame the user is still looking at is the last frame + // painted from defaults. + useEffect(() => { + const restoredUi = readUiState(); + const restoredTabs = readWorkspaceTabs(); + setPersisted(restoredUi); + setWorkspaceTabs(restoredTabs); + setTabState(restoredTabs.tabStrip); + setColumnState(restoredTabs.columnLayout); + setPanel(restoredUi.panel); + setUiRestored(true); + }, []); // Mirror panel changes into localStorage. The write helper is // debounced; mounting/de-mounting the panel quickly during a - // refresh never floods storage. + // refresh never floods storage. Gated on `uiRestored` so the + // defaults-seeded first render cannot clobber the stored payload + // (webui-parity 106). useEffect(() => { + if (!uiRestored) return; writeUiState({ ...DEFAULT_UI_STATE, ...persisted, @@ -129,15 +163,23 @@ function App() { // intentionally not adding `persisted` to deps — the persist // module already guards the debounced write. // eslint-disable-next-line react-hooks/exhaustive-deps - }, [panel]); + }, [panel, uiRestored]); // Slice 17 — mirror workspace-tabs state into localStorage. The // debounced writer coalesces open + activate + scroll edits into - // one write. + // one write. Gated on `uiRestored` for the same reason as above. useEffect(() => { + if (!uiRestored) return; writeWorkspaceTabs({ tabStrip: tabState, columnLayout: columnState }); // eslint-disable-next-line react-hooks/exhaustive-deps - }, [tabState, columnState]); + }, [tabState, columnState, uiRestored]); + + // Slice 17 — WorkspaceColumns self-measures via + // ResizeObserver, so the page does not need to feed + // containerWidth anymore. viewportWidth is still threaded + // through for the future auto-collapse ladder; today it is + // accepted but unused inside computeColumnLayout. + const viewportWidth = typeof window !== "undefined" ? window.innerWidth : 1280; // ============================================================ // Workspace tabs reducers (the page wires every action through @@ -477,6 +519,12 @@ function App() { useEffect(() => { if (!urlRestored) return; + // webui-parity 106 — same gate as the other write mirrors: this effect + // can fire before the mount restore has applied the stored payload + // (SSE sometimes names a session before the restore batch lands), and + // writing from the defaults-seeded `persisted` would drop the stored + // appearance / sidebar choice. + if (!uiRestored) return; const active = state?.mcodeSessionId ?? null; writeUiState({ ...DEFAULT_UI_STATE, @@ -485,7 +533,7 @@ function App() { lastSessionId: active, }); // eslint-disable-next-line react-hooks/exhaustive-deps - }, [state?.mcodeSessionId, urlRestored]); + }, [state?.mcodeSessionId, urlRestored, uiRestored]); useEffect(() => { const onPop = () => { @@ -739,7 +787,16 @@ function App() { } /** - * Scroll-restored chat wrapper — unchanged from slice 07. + * Scroll-restored chat wrapper. + * + * webui-parity 106: the page no longer reads the scroll position out of + * localStorage during render (the pre-106 render-phase read was the third + * instance of the hydration bomb). `Chat` already re-reads the SAME + * per-session key (`webui:scroll:v1::`) inside its own + * post-mount restore effect whenever `sessionKey` changes, and it falls back + * to that read whenever no explicit scroll target arrives — so passing + * nothing restores the identical number, from an effect that only runs + * client-side. The page keeps only the write half of the contract. */ function ScrollRestoredChat({ t, @@ -752,13 +809,11 @@ function ScrollRestoredChat({ sessionId: string | null; onOpenFile?: (path: string) => void; }) { - const initial = sessionId ? readScrollPosition(sessionId) : 0; return ( { if (!sessionId) return; writeScrollPosition(sessionId, top); @@ -769,10 +824,9 @@ function ScrollRestoredChat({ } // keep the unused-export lint happy: slice 17 deliberately does -// not pull `panel` / `openPanel` / `openSettings` / `DEFAULT_UI_STATE` -// from the legacy path. They stay in scope so a future ticket can -// revive them without re-importing the modules. +// not pull `panel` / `openPanel` / `openSettings` from the legacy +// path. They stay in scope so a future ticket can revive them +// without re-importing the modules. void PreviewColumn; -void DEFAULT_WORKSPACE_TABS_STATE; void isHtmlPath; void useRef; \ No newline at end of file diff --git a/packages/webui/webapp/components/composer.tsx b/packages/webui/webapp/components/composer.tsx index afb9dbca..623ad3d8 100644 --- a/packages/webui/webapp/components/composer.tsx +++ b/packages/webui/webapp/components/composer.tsx @@ -36,6 +36,7 @@ import { mergeRestoredDraft, setComposerDraft, subscribeComposerDraft, + unconfirmedPatchOnTurnEnd, } from "@/lib/composer-draft"; import { completeComposerSent, @@ -159,21 +160,38 @@ export function Composer({ onAddProvider?: () => void; }) { const { state, providersRevision } = useSessionContext(); + // The session this composer is standing in. Derived before the draft + // subscription because the draft store is keyed BY SESSION (webui-parity + // 106, smoke-report P5): the getter below reads this session's box, so a + // session switch swaps text, attachments and the banner synchronously in + // the same render instead of bleeding the previous session's state in. + // `""` is the no-session bucket (home screen, before the first snapshot). + const modelKey = state?.model?.name ?? ""; + const sessionKey = state?.sessionId ?? ""; // Text, attachments, and the error banner live in the module-scope draft // store (lib/composer-draft.ts) rather than useState: page.tsx swaps this // component between two tree positions when the first conversation line // lands in a state push, and a `useState`-held draft died with the // unmounted instance. The store survives the swap, so whatever the user // typed — and the failure banner they need to read — outlives any - // remount. `sending` stays local: it is per-submit bookkeeping, not user - // input worth preserving. - const draft = useSyncExternalStore(subscribeComposerDraft, getComposerDraft, getComposerDraft); + // remount. Since 106 the store is per-session: the keyed getter keeps + // session A's draft out of session B's composer, and both drafts survive + // the round trip. `sending` stays local: it is per-submit bookkeeping, not + // user input worth preserving. + const draft = useSyncExternalStore( + subscribeComposerDraft, + () => getComposerDraft(sessionKey), + () => getComposerDraft(""), + ); const value = draft.value; const attachments = draft.attachments; const error = draft.error; const errorKind = draft.errorKind; const unconfirmedOutcome = draft.unconfirmed; - const setValue = useCallback((next: string) => setComposerDraft({ value: next }), []); + const setValue = useCallback( + (next: string) => setComposerDraft(sessionKey, { value: next }), + [sessionKey], + ); const [sending, setSending] = useState(false); const [models, setModels] = useState< { @@ -240,6 +258,26 @@ export function Composer({ /** Nothing to send yet — the send button is rendered but inert. */ const empty = value.trim().length === 0 && attachments.length === 0; + // Smoke-report P4 (webui-parity 106): the grey unconfirmed banner must not + // outlive the turn it warned about. When the acknowledgement timed out, the + // probe answered "the engine is running this message — do not resend"; once + // `running.active` falls, that warning describes a turn that is over, and + // after `sleep 35` it used to sit under the input until the next send or a + // reload. The decision lives in `unconfirmedPatchOnTurnEnd` (unit-tested); + // the wiring here only feeds it the running-flag fall. #126's three-value + // display semantics are untouched — this owns dismissal, not display, and + // a real `rejected` refusal keeps its dismiss paths. + const prevRunningRef = useRef(running); + useEffect(() => { + const patch = unconfirmedPatchOnTurnEnd( + prevRunningRef.current, + running, + errorKind, + ); + prevRunningRef.current = running; + if (patch) setComposerDraft(sessionKey, patch); + }, [running, errorKind, sessionKey]); + // The model catalogue comes from the server; the chip shows the active model // from the state snapshot so it tracks changes made elsewhere. // @@ -251,18 +289,10 @@ export function Composer({ // binary or a saved models.json takes effect on the next chip open. We // re-fetch when the session or the active model changes rather than only // on mount. - const modelKey = state?.model?.name ?? ""; - const sessionKey = state?.sessionId ?? ""; - // The failure banner is scoped to the session it failed in. `composer-draft` - // is a module-scope store shared by every composer instance (it has to be — - // page.tsx swaps the composer between two tree positions), so without this - // reset a rejection recorded in session A rode along when the user switched - // to session B and painted B's composer red for a send B never made. The - // typed draft is deliberately NOT cleared: the user's words belong to them, - // and the restored-draft merge below already owns cross-session text rules. - useEffect(() => { - setComposerDraft({ error: null, errorKind: null, unconfirmed: null }); - }, [sessionKey]); + // The banner no longer needs a sessionKey-keyed clear effect: since 106 the + // draft store itself is keyed by session, so a banner recorded in session A + // simply lives in A's box and session B reads its own (empty) one. The + // typed draft stays with its session for the same reason. useEffect(() => { void api .listModels() @@ -404,10 +434,20 @@ export function Composer({ // The outbox record stores these, so a later failure can identify // its owner. They are NOT the values the catch branch compares // against; the catch branch reads the LIVE context (see below). + // `dispatchDraftKey` is the same identity in the per-session draft + // store: the banner a failed send leaves behind must land in the + // session that attempted it, so the user finds it when they come + // back — never pasted into whichever session they are looking at + // by then (smoke-report P5, the s28 capture). const dispatchCid = clientId(); const dispatchSessionId = state?.sessionId ?? null; + const dispatchDraftKey = dispatchSessionId ?? ""; setSending(true); - setComposerDraft({ error: null, errorKind: null, unconfirmed: null }); + setComposerDraft(dispatchDraftKey, { + error: null, + errorKind: null, + unconfirmed: null, + }); // Ticket 13 — optimistic clear. The backend does session // switching and transcript backfill before its ack, so waiting // for the await leaves the text sitting in the box for the whole @@ -424,7 +464,7 @@ export function Composer({ content, attachments, }); - setComposerDraft({ value: "", attachments: [] }); + setComposerDraft(dispatchDraftKey, { value: "", attachments: [] }); try { // A slash input is a message OR a command, and only the eight // webui button commands belong to /api/cmd — routing on the @@ -481,16 +521,22 @@ export function Composer({ // failComposerSent returns the restore payload only when the // LIVE context still matches the dispatch context — a session // switch mid-flight must never paste the old session's text - // into the new session's composer. + // into the new session's composer. When it does match, the live + // session IS the dispatch session, so keying the merge by the + // live draft key writes the same box the user is looking at. const restored = failComposerSent({ cid: liveCid, sessionId: liveSessionId, error: errorMessage, }); - // Always set the error banner — the failure is real even when - // the active session no longer matches the record (the banner - // is in the module-scope draft store too, so it outlives a - // session switch). + // Always set the error banner — but in the DISPATCH session's + // draft box, not the live one. The failure is real even when the + // active session no longer matches the record; with the per- + // session store, writing it into the owning session means the + // user finds the banner when they return to that session, and + // the session they switched TO never paints red for a send it + // never made (the s28 bleed in the smoke report). The submit- + // path clear above already keyed the same box. if (restored && (outcome === null || shouldRestoreDraft(outcome))) { // A send the server may already be running must NOT come back as text // sitting in the box: one Enter would run it a second time. The @@ -501,9 +547,12 @@ export function Composer({ // whatever was typed since. The merge rule lives in // lib/composer-draft#mergeRestoredDraft so it is unit-tested // instead of being re-derived from a React callback. - setComposerDraft(mergeRestoredDraft(getComposerDraft(), restored)); + setComposerDraft( + dispatchDraftKey, + mergeRestoredDraft(getComposerDraft(dispatchDraftKey), restored), + ); } - setComposerDraft({ + setComposerDraft(dispatchDraftKey, { error: errorMessage, errorKind: unconfirmed ? "unconfirmed" : "rejected", unconfirmed: outcome, @@ -521,13 +570,17 @@ export function Composer({ const result = await api.uploadFile(file); if (result?.path) picked.push(`@${result.path}`); } catch (cause) { - setComposerDraft({ error: cause instanceof Error ? cause.message : String(cause) }); + setComposerDraft(sessionKey, { + error: cause instanceof Error ? cause.message : String(cause), + }); } } if (picked.length) { - setComposerDraft((current) => ({ attachments: [...current.attachments, ...picked] })); + setComposerDraft(sessionKey, (current) => ({ + attachments: [...current.attachments, ...picked], + })); } - }, []); + }, [sessionKey]); // Drag-and-drop file upload. `preventDefault` on `dragover` is required: // without it the browser opens the file in the tab. Text drags (selecting @@ -764,6 +817,7 @@ export function Composer({ groups={groups} value={state?.model.name} label={currentModelLabel} + sessionKey={sessionKey} thinking={state?.model?.thinking ?? ""} contextWindow={currentContextWindow} onAddProvider={onAddProvider} @@ -1127,6 +1181,7 @@ function ModelSelect({ groups, value, label, + sessionKey, thinking, contextWindow, onPick, @@ -1136,6 +1191,13 @@ function ModelSelect({ onAddProvider, }: { t: (key: MessageKey) => string; + /** The active session id. The picker's local UI state (open cascade, + * previewed row, per-model draft mirror) is session-scoped bookkeeping: + * on a session switch it resets, so no session-A menu state visually + * persists into session B's view (webui-parity 106, smoke-report P5). + * The chip VALUE is not local — it reads the server's state snapshot — + * so per-session model truth rides the same SSE path as before. */ + sessionKey: string; models: { id: string; label: string; @@ -1218,6 +1280,18 @@ function ModelSelect({ const [drafts, setDrafts] = useState< Record >({}); + // webui-parity 106 — the four local states above belong to ONE session's + // picker interaction. A session switch that arrived while the cascade was + // open (or a preview row focused) used to carry all of it into the next + // session's view. The reset is a no-op while the session is stable — the + // effect only fires on a real key change, and closing an already-closed + // cascade writes nothing. + useEffect(() => { + setOpen(false); + setSubmenuFor(null); + setFocusedModelId(null); + setDrafts({}); + }, [sessionKey]); /** Ref to the provider row that owns the open submenu. */ const submenuAnchorRef = useRef(null); /** Ref to the provider row that contains the active model, so the diff --git a/packages/webui/webapp/lib/composer-draft.ts b/packages/webui/webapp/lib/composer-draft.ts index 3e98db70..6fa658d4 100644 --- a/packages/webui/webapp/lib/composer-draft.ts +++ b/packages/webui/webapp/lib/composer-draft.ts @@ -1,6 +1,6 @@ /** * Composer draft store — the typed text, the `@path` attachment chips, and - * the send-error banner, held OUTSIDE the React tree. + * the send-error banner, held OUTSIDE the React tree and keyed BY SESSION. * * Why module scope instead of component state: `app/page.tsx` swaps the * composer between two tree positions — `(); const listeners = new Set<() => void>(); export type ComposerDraftPatch = | Partial | ((current: ComposerDraft) => Partial); -/** Write a patch (or an updater, mirroring `setState` semantics). */ -export function setComposerDraft(patch: ComposerDraftPatch): void { - const resolved = typeof patch === "function" ? patch(draft) : patch; - draft = { ...draft, ...resolved }; - for (const listener of listeners) listener(); +/** + * Read the draft of ONE session. Stable identity between writes — the + * returned object only changes when that session's draft is written, and + * the shared `EMPTY_DRAFT` singleton stands in for sessions without one, + * so `useSyncExternalStore` can compare by reference. + */ +export function getComposerDraft(sessionKey: string): ComposerDraft { + return drafts.get(sessionKey) ?? EMPTY_DRAFT; } -/** Read the current draft. Stable identity between writes. */ -export function getComposerDraft(): ComposerDraft { - return draft; +/** Write a patch (or an updater, mirroring `setState` semantics) into ONE + * session's draft. Other sessions' drafts are untouched — that is the + * isolation contract the smoke report's P5 depends on. */ +export function setComposerDraft( + sessionKey: string, + patch: ComposerDraftPatch, +): void { + const current = drafts.get(sessionKey) ?? EMPTY_DRAFT; + const resolved = typeof patch === "function" ? patch(current) : patch; + drafts.set(sessionKey, { ...current, ...resolved }); + for (const listener of listeners) listener(); } /** `useSyncExternalStore` subscription. Returns the unsubscribe thunk. */ @@ -100,12 +128,6 @@ export function subscribeComposerDraft(listener: () => void): () => void { }; } -/** Test-only: reset the draft to empty between cases. */ -export function resetComposerDraftForTests(): void { - draft = EMPTY_DRAFT; - listeners.clear(); -} - /** * The patch that puts a rejected submission back into the composer * without clobbering what the user typed while it was in flight. @@ -140,3 +162,37 @@ export function mergeRestoredDraft( attachments: [...restored.attachments, ...current.attachments], }; } + +/** + * The patch to apply when a turn ends, or `null` for "nothing to do". + * + * The unconfirmed banner's three-value display semantics are #126's and are + * NOT touched here — this only owns WHEN the banner goes away. A send whose + * acknowledgement timed out leaves a grey "the engine is running this + * message — do not resend" banner; once the turn it warned about is over, + * the warning describes nothing and must disappear (smoke-report P4: after + * `sleep 35` completed, the banner stayed until the next send or reload). + * A real `rejected` refusal is a different fact and stays until the user + * acts on it. + * + * The turn-end signal is the running flag falling: `prevRunning === true` + * and `running === false`. A banner that appears while no turn runs (the + * fast-turn echo path) never sees that fall inside the same mount, so it + * keeps the pre-existing dismiss paths — the next send in the same session + * clears it, as does a session switch (per-session isolation, above). + */ +export function unconfirmedPatchOnTurnEnd( + prevRunning: boolean, + running: boolean, + errorKind: ComposerErrorKind | null, +): ComposerDraftPatch | null { + if (!(prevRunning && !running)) return null; + if (errorKind !== "unconfirmed") return null; + return { error: null, errorKind: null, unconfirmed: null }; +} + +/** Test-only: reset every session's draft and the listeners between cases. */ +export function resetComposerDraftForTests(): void { + drafts.clear(); + listeners.clear(); +} diff --git a/packages/webui/webapp/test/composer-draft.test.ts b/packages/webui/webapp/test/composer-draft.test.ts index 309133e4..e47f448f 100644 --- a/packages/webui/webapp/test/composer-draft.test.ts +++ b/packages/webui/webapp/test/composer-draft.test.ts @@ -19,6 +19,15 @@ // 3. Subscribers are notified on writes and stopped by unsubscribe. // 4. The updater form works against the CURRENT draft (no stale // closure over an older snapshot). +// +// webui-parity 106 (smoke-report P5) adds the isolation contract: the +// store is keyed BY SESSION, so session A's draft — text, chips, banner — +// is invisible to session B and still there when the user comes back. The +// smoke run's s28 capture (session 2's view showing session 1's draft, +// GLM chip and 409 banner) is the regression these tests fence off. The +// same ticket's P4 fix pins `unconfirmedPatchOnTurnEnd`, the pure decision +// that retires the grey unconfirmed banner when the turn it warned about +// ends. import { test, describe, beforeEach } from "node:test"; import assert from "node:assert/strict"; @@ -28,36 +37,41 @@ import { resetComposerDraftForTests, setComposerDraft, subscribeComposerDraft, + unconfirmedPatchOnTurnEnd, + type ComposerDraft, } from "../lib/composer-draft"; beforeEach(() => { resetComposerDraftForTests(); }); +const S1 = "mvs_session_one"; +const S2 = "mvs_session_two"; + describe("composer draft store — survives the composer's remount", () => { test("typed text written before the 'swap' is read by a fresh reader after it", () => { // The composer instance that existed before the remount wrote the text. - setComposerDraft({ value: "继续这个任务" }); + setComposerDraft(S1, { value: "继续这个任务" }); // A fresh mount reads the module-scope store — the same object any // later instance sees, regardless of tree position. - const afterRemount = getComposerDraft(); + const afterRemount = getComposerDraft(S1); assert.equal(afterRemount.value, "继续这个任务"); }); test("send-error banner survives the same remount", () => { // A 409 session-busy failure wrote the banner right before the push // that remounts the composer. - setComposerDraft({ error: "this conversation is already running in another window" }); - assert.equal(getComposerDraft().error, "this conversation is already running in another window"); + setComposerDraft(S1, { error: "this conversation is already running in another window" }); + assert.equal(getComposerDraft(S1).error, "this conversation is already running in another window"); }); test("attachment chips survive the remount too", () => { - setComposerDraft({ attachments: ["@uploads/a.txt"] }); - assert.deepEqual(getComposerDraft().attachments, ["@uploads/a.txt"]); + setComposerDraft(S1, { attachments: ["@uploads/a.txt"] }); + assert.deepEqual(getComposerDraft(S1).attachments, ["@uploads/a.txt"]); }); test("nothing resets the draft implicitly — only explicit writes do", () => { - setComposerDraft({ + setComposerDraft(S1, { value: "draft", error: "boom", attachments: ["@uploads/a.txt"], @@ -66,16 +80,80 @@ describe("composer draft store — survives the composer's remount", () => { // the store never mutates it, and there is no reset hook on the // production surface. for (let i = 0; i < 3; i++) { - assert.equal(getComposerDraft().value, "draft"); - assert.equal(getComposerDraft().error, "boom"); + assert.equal(getComposerDraft(S1).value, "draft"); + assert.equal(getComposerDraft(S1).error, "boom"); } // The only clearing write is the composer's own success path. - setComposerDraft({ value: "", attachments: [] }); - assert.equal(getComposerDraft().value, ""); - assert.deepEqual(getComposerDraft().attachments, []); + setComposerDraft(S1, { value: "", attachments: [] }); + assert.equal(getComposerDraft(S1).value, ""); + assert.deepEqual(getComposerDraft(S1).attachments, []); // The error banner persists until the next submit start clears it — // matches the pre-existing `setError(null)`-at-submit semantics. - assert.equal(getComposerDraft().error, "boom"); + assert.equal(getComposerDraft(S1).error, "boom"); + }); +}); + +describe("composer draft store — per-session isolation (webui-parity 106)", () => { + test("session B never sees session A's typed draft", () => { + setComposerDraft(S1, { value: "计时10s,后说hi" }); + assert.equal(getComposerDraft(S2).value, ""); + assert.equal(getComposerDraft(S2).error, null); + assert.deepEqual(getComposerDraft(S2).attachments, []); + }); + + test("switching back restores session A's draft untouched", () => { + setComposerDraft(S1, { value: "A 的草稿" }); + // The user works in B for a while — types, fails a send, clears. + setComposerDraft(S2, { value: "B 的草稿" }); + setComposerDraft(S2, { error: "HTTP 500" }); + setComposerDraft(S2, { value: "", attachments: [] }); + // Back to A: the round trip must be lossless. + assert.equal(getComposerDraft(S1).value, "A 的草稿"); + assert.equal(getComposerDraft(S1).error, null); + }); + + test("a banner written to its owning session does not paint the other one", () => { + // The catch branch writes the banner into the DISPATCH session's box + // even though the user is already looking at another session. + setComposerDraft(S1, { error: "a turn is already running for this client", errorKind: "rejected" }); + assert.equal(getComposerDraft(S2).error, null); + assert.equal(getComposerDraft(S2).errorKind, null); + // Returning to A still shows it — the failure belongs to A. + assert.equal(getComposerDraft(S1).error, "a turn is already running for this client"); + }); + + test("an unread banner in one session survives a visit to the other", () => { + setComposerDraft(S1, { error: "boom", errorKind: "rejected" }); + setComposerDraft(S2, { value: "unrelated work" }); + assert.equal(getComposerDraft(S1).error, "boom", "A's banner must not be cleared by visiting B"); + assert.equal(getComposerDraft(S2).value, "unrelated work"); + }); + + test("the no-session bucket is its own box", () => { + // The home screen (before the first snapshot names a session) reads + // the "" key; its draft must not bleed into a real session either. + setComposerDraft("", { value: "home draft" }); + assert.equal(getComposerDraft(S1).value, ""); + assert.equal(getComposerDraft("").value, "home draft"); + }); + + test("the empty-session read is a stable singleton until written", () => { + // `useSyncExternalStore` compares snapshots by reference; an unstable + // identity for missing drafts would loop the subscription. + assert.equal(getComposerDraft("never-written"), getComposerDraft("also-never-written")); + const before = getComposerDraft(S1); + setComposerDraft(S2, { value: "b" }); + assert.equal(getComposerDraft(S1), before, "writing B must not change A's snapshot identity"); + }); + + test("updater form patches against the CURRENT session's draft", () => { + setComposerDraft(S1, { attachments: ["@one"] }); + setComposerDraft(S2, { attachments: ["@b-one"] }); + setComposerDraft(S2, (current: ComposerDraft) => ({ + attachments: [...current.attachments, "@b-two"], + })); + assert.deepEqual(getComposerDraft(S2).attachments, ["@b-one", "@b-two"]); + assert.deepEqual(getComposerDraft(S1).attachments, ["@one"], "A untouched by B's updater"); }); }); @@ -83,25 +161,72 @@ describe("composer draft store — subscription", () => { test("listeners are notified on every write", () => { const seen: string[] = []; const unsubscribe = subscribeComposerDraft(() => { - seen.push(getComposerDraft().value); + seen.push(getComposerDraft(S1).value); }); - setComposerDraft({ value: "a" }); - setComposerDraft({ value: "ab" }); + setComposerDraft(S1, { value: "a" }); + setComposerDraft(S1, { value: "ab" }); unsubscribe(); - setComposerDraft({ value: "abc" }); + setComposerDraft(S1, { value: "abc" }); assert.deepEqual(seen, ["a", "ab"]); }); - test("updater form patches against the CURRENT draft", () => { - setComposerDraft({ attachments: ["@one"] }); - setComposerDraft((current) => ({ - attachments: [...current.attachments, "@two"], - })); - assert.deepEqual(getComposerDraft().attachments, ["@one", "@two"]); - // Patches merge — an attachments write must not drop `value`. - setComposerDraft({ value: "text" }); - setComposerDraft((current) => ({ attachments: [...current.attachments, "@three"] })); - assert.equal(getComposerDraft().value, "text"); - assert.deepEqual(getComposerDraft().attachments, ["@one", "@two", "@three"]); + test("a write to ANY session notifies subscribers (the composer re-checks its key)", () => { + // The keyed `useSyncExternalStore` getter re-reads on notification; a + // write that no listener ever hears about could leave a stale box on + // screen after a switch. + let notified = 0; + const unsubscribe = subscribeComposerDraft(() => { + notified += 1; + }); + setComposerDraft(S1, { value: "a" }); + setComposerDraft(S2, { value: "b" }); + unsubscribe(); + assert.equal(notified, 2); + }); + + test("patches merge — an attachments write must not drop `value`", () => { + setComposerDraft(S1, { value: "text" }); + setComposerDraft(S1, { attachments: ["@one"] }); + assert.equal(getComposerDraft(S1).value, "text"); + assert.deepEqual(getComposerDraft(S1).attachments, ["@one"]); + }); +}); + +describe("unconfirmedPatchOnTurnEnd — the grey banner dies with its turn (P4)", () => { + const unconfirmed: Pick = { errorKind: "unconfirmed" }; + + test("running falling (true → false) clears the unconfirmed banner", () => { + // The smoke-report scenario: `sleep 35` timed out at the 30s ack + // deadline, the probe said "accepted", the turn finished — and the + // grey banner stayed on screen. The fall is the retire signal. + assert.deepEqual( + unconfirmedPatchOnTurnEnd(true, false, unconfirmed.errorKind), + { error: null, errorKind: null, unconfirmed: null }, + ); + }); + + test("a turn still running keeps the banner", () => { + assert.equal(unconfirmedPatchOnTurnEnd(true, true, "unconfirmed"), null); + }); + + test("no observed turn (false → false, e.g. banner set after a fast turn) keeps it", () => { + // The fast-turn echo path never shows a running fall inside the same + // mount; the pre-existing dismiss paths (next send, session switch) + // stay responsible for that case. #126's display semantics untouched. + assert.equal(unconfirmedPatchOnTurnEnd(false, false, "unconfirmed"), null); + }); + + test("a turn starting (false → true) never clears anything", () => { + assert.equal(unconfirmedPatchOnTurnEnd(false, true, "unconfirmed"), null); + }); + + test("a real rejected refusal is NOT cleared by the turn ending", () => { + // "消息发送失败" is a different fact — it stays until the user acts on + // it or the next send in that session starts. + assert.equal(unconfirmedPatchOnTurnEnd(true, false, "rejected"), null); + }); + + test("no banner at all → no patch", () => { + assert.equal(unconfirmedPatchOnTurnEnd(true, false, null), null); }); }); diff --git a/packages/webui/webapp/test/composer-submit-tripwire.test.ts b/packages/webui/webapp/test/composer-submit-tripwire.test.ts index 7ca9569b..80c47f12 100644 --- a/packages/webui/webapp/test/composer-submit-tripwire.test.ts +++ b/packages/webui/webapp/test/composer-submit-tripwire.test.ts @@ -82,8 +82,8 @@ describe("composer submit ordering — ticket 13 wiring tripwire", () => { ); const clearDraftIdx = indexOfOrThrow( composerSource, - 'setComposerDraft({ value: "", attachments: [] })', - 'setComposerDraft({ value: "", attachments: [] })', + 'setComposerDraft(dispatchDraftKey, { value: "", attachments: [] })', + 'setComposerDraft(dispatchDraftKey, { value: "", attachments: [] })', ); const awaitSendIdx = indexOfOrThrow( composerSource, @@ -246,38 +246,161 @@ describe("composer submit ordering — ticket 13 wiring tripwire", () => { ); }); }); -describe("switching sessions clears the failure banner", () => { - // The error banner lives in the module-scope composer-draft store, which is - // shared by every composer instance (page.tsx swaps the composer between two - // tree positions, so a useState-held draft would die on the swap). Without a - // per-session reset, a rejection recorded in session A kept painting - // session B's composer red after the switch — a send B never made, with a - // red "消息发送失败: HTTP 500" banner appearing "on switching" (P5/P6 of the - // 100-ticket smoke report). The draft TEXT is deliberately not cleared: the - // restored-draft merge owns cross-session text rules. - test("an effect keyed on sessionKey resets error/errorKind/unconfirmed", () => { - // The effect body and the submit-path reset must carry the same three - // fields — a banner kind added later has to join both, and a revert that - // drops the effect (or re-keys it to something that never changes, like a - // stable ref) fails the dependency-array assertion. +describe("the draft store is keyed by session (webui-parity 106, smoke P5)", () => { + // The pre-106 store was ONE shared bucket: session A's draft, chips and + // failure banner rode into session B's view on a switch (the s28 capture + // in the smoke report). #141 papered over the banner half with a + // sessionKey-keyed clear effect — which also destroyed the banner of the + // session the user was RETURNING to. Since 106 the isolation is + // structural: the composer reads and writes the store THROUGH the active + // session key, and the catch branch writes the banner into the DISPATCH + // session's box. These tripwires pin that wiring; the store-level + // behaviour itself is unit-tested in composer-draft.test.ts. + + test("the composer's draft snapshot is read through the session key", () => { + // A revert to the shared bucket re-appears as `getComposerDraft` being + // called with NO key in the useSyncExternalStore call. assert.match( composerSource, - /useEffect\(\(\) => \{\s*setComposerDraft\(\{\s*error: null,\s*errorKind: null,\s*unconfirmed: null,?\s*\}\);\s*\}, \[sessionKey\]\);/, - "composer must reset the banner fields in an effect keyed on sessionKey — " + - "the module-scope draft store outlives sessions, so the banner must be " + - "scoped to the session it failed in", + /useSyncExternalStore\(\s*subscribeComposerDraft,\s*\(\) => getComposerDraft\(sessionKey\),\s*\(\) => getComposerDraft\(""\),?\s*\)/, + "the draft snapshot must be read through the session key so a switch " + + "swaps boxes synchronously — a key-less getter is the shared-bucket " + + "regression this ticket fixes", + ); + }); + + test("the #141 clear effect is gone — isolation is structural now", () => { + // The clear-on-switch effect destroyed a RETURNING session's own + // unread banner. With per-session boxes it is wrong in every case. + assert.ok( + !/useEffect\(\(\) => \{\s*setComposerDraft\(\{?\s*error: null,\s*errorKind: null,\s*unconfirmed: null,?\s*\}?\);?\s*\}, \[sessionKey\]\);/.test( + composerSource, + ), + "the sessionKey-keyed banner-clear effect must not come back — the " + + "keyed store already isolates banners per session, and clearing on " + + "switch loses the session the user returns to", + ); + }); + + test("every submit-path write carries a session key", () => { + // A key-less setComposerDraft call site would write into whatever + // box... nothing — it is a type error; the tripwire pins the two + // load-bearing literals so a refactor that drops the key from them + // fails here rather than silently changing boxes. + assert.ok( + /setComposerDraft\(dispatchDraftKey, \{\s*error: null,\s*errorKind: null,\s*unconfirmed: null,?\s*\}\)/.test( + composerSource, + ), + "submit must clear the banner in the DISPATCH session's box", + ); + assert.match( + composerSource, + /setComposerDraft\(dispatchDraftKey, \{ value: "", attachments: \[\] \}\)/, + "the optimistic clear must write the dispatch session's box", + ); + }); + + test("the catch branch writes the banner into the dispatch session's box", () => { + // The failure belongs to the session that attempted the send. Writing + // it into the LIVE key would repaint the session the user switched TO + // — the exact bleed the s28 capture shows. + const catchStartIdx = indexOfOrThrow( + composerSource, + "} catch (cause) {", + "} catch (cause) {", + ); + const catchBody = composerSource.slice(catchStartIdx); + assert.match( + catchBody, + /setComposerDraft\(dispatchDraftKey, \{\s*error: errorMessage,/, + "the banner must be keyed by dispatchDraftKey inside the catch branch", + ); + assert.ok( + !/setComposerDraft\(sessionKey, \{[^}]*errorMessage/.test(catchBody), + "the banner must NOT be written into the live session's box — a send " + + "that failed in session A must never paint session B red", ); }); test("the submit path still clears the banner before dispatching", () => { - // The session-switch reset is additive; it must not replace the - // clear-on-submit (a retry in the SAME session also has to clear the old - // rejection before the new attempt is judged). + // A same-session retry also has to clear the old rejection before the + // new attempt is judged. assert.match( composerSource, - /setSending\(true\);\s*setComposerDraft\(\{\s*error: null,\s*errorKind: null,\s*unconfirmed: null,?\s*\}\);/, + /setSending\(true\);\s*setComposerDraft\(dispatchDraftKey, \{\s*error: null,\s*errorKind: null,\s*unconfirmed: null,?\s*\}\);/, "submit must clear the banner right after setSending(true), before the " + "optimistic park — a same-session retry starts clean", ); }); }); + +describe("the unconfirmed banner retires when its turn ends (webui-parity 106, smoke P4)", () => { + // After `sleep 35` finished, the grey "服务器一直没有确认" banner stayed + // under the input until the next send or a reload. The decision + // (running-flag fall + errorKind === "unconfirmed") is unit-tested in + // composer-draft.test.ts; this pins the WIRING: the composer must feed + // the turn-end transition into it and apply the patch it returns. + + test("the composer watches the running flag and applies the turn-end patch", () => { + assert.match( + composerSource, + /const prevRunningRef = useRef\(running\);/, + "the previous running value must be captured per render", + ); + const effectIdx = indexOfOrThrow( + composerSource, + "unconfirmedPatchOnTurnEnd(", + "unconfirmedPatchOnTurnEnd( call", + ); + const wiring = composerSource.slice(effectIdx - 200, effectIdx + 400); + assert.match( + wiring, + /prevRunningRef\.current,\s*running,\s*errorKind,/, + "the decision must receive (previous running, running, errorKind)", + ); + assert.match( + wiring, + /prevRunningRef\.current = running;/, + "the reference must advance after the decision, or one stale value " + + "would clear (or keep) the banner on unrelated re-renders", + ); + assert.match( + wiring, + /if \(patch\) setComposerDraft\(sessionKey, patch\);/, + "a non-null patch must be applied to the ACTIVE session's box", + ); + }); + + test("the turn-end decision is imported from the draft module", () => { + assert.match( + composerSource, + /import\s+\{[^}]*\bunconfirmedPatchOnTurnEnd\b[^}]*\}\s+from\s+["']@\/lib\/composer-draft["']/, + "the decision must be the product function, not an inline re-derivation", + ); + }); +}); + +describe("the model picker's local state resets on a session switch (webui-parity 106, smoke P5)", () => { + // The chip VALUE reads the server snapshot, but the cascade's open flag, + // previewed row and per-model draft mirror are component-local; without a + // reset, session A's open menu / preview state visually persisted into + // session B's view. + + test("ModelSelect receives the session key", () => { + assert.match( + composerSource, + /]*sessionKey=\{sessionKey\}/, + "the composer must pass the session key down to the picker", + ); + }); + + test("the picker resets its local states in an effect keyed on sessionKey", () => { + const resetIdx = indexOfOrThrow( + composerSource, + "setOpen(false);\n setSubmenuFor(null);\n setFocusedModelId(null);\n setDrafts({});", + "ModelSelect's four local-state resets", + ); + const deps = composerSource.slice(resetIdx, resetIdx + 200); + assert.match(deps, /\}, \[sessionKey\]\);/, "the reset must be keyed on sessionKey"); + }); +}); diff --git a/packages/webui/webapp/test/page-hydration.test.ts b/packages/webui/webapp/test/page-hydration.test.ts new file mode 100644 index 00000000..4fd7ca35 --- /dev/null +++ b/packages/webui/webapp/test/page-hydration.test.ts @@ -0,0 +1,153 @@ +// webapp/test/page-hydration.test.ts +// +// Static-source tripwires for the page root's storage-access timing +// (webui-parity 106, smoke-report P7-b) and for the scroll-restore +// contract that must survive it (red line: 刷新后滚动位置还在). +// +// Why a tripwire and not a unit test: this suite has no React render +// harness (plain `node --test` over the lib modules), and the defect is +// not a function's output but WHERE a function is called from — the +// render phase of the prerendered root component. `app/page.tsx` is +// pre-rendered by the Next.js static export, so the server HTML and the +// client's first (hydration) render must be byte-identical. Any +// `localStorage` read that runs during render returns defaults on the +// server and stored values on the client — a hydration mismatch that +// today is masked by the `state === null` skeleton and detonates the +// moment that skeleton changes. The reads therefore live in exactly one +// place: the post-mount restore effect. +// +// The load-bearing survivor of that move is the transcript scroll +// restore: the page dropped its render-phase `readScrollPosition` call +// because `Chat` already re-reads the SAME per-session key inside its own +// post-mount effect (and falls back to it whenever `initialScrollTop` is +// absent). That fallback IS the restore behaviour now, so it is pinned +// here too. + +import { test, describe } from "node:test"; +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import { fileURLToPath } from "node:url"; +import { resolve, dirname } from "node:path"; + +const here = dirname(fileURLToPath(import.meta.url)); +const pageSource = readFileSync(resolve(here, "../app/page.tsx"), "utf8"); +const chatSource = readFileSync(resolve(here, "../components/chat.tsx"), "utf8"); + +describe("page.tsx never reads storage during render (webui-parity 106)", () => { + test("no useState initializer calls a persist reader", () => { + // The pre-106 shape — `useState(() => readUiState())` — ran + // localStorage reads on every (re)render entry, server included. + assert.ok( + !/useState[^;]*\(\)\s*=>\s*read(?:UiState|WorkspaceTabs|ScrollPosition)\(/.test(pageSource), + "page.tsx must not seed state from a storage read — the initializer " + + "runs during render, and render runs on the static-export server " + + "too (the hydration bomb P7-b)", + ); + }); + + test("readScrollPosition is gone from the page entirely", () => { + // The scroll wrapper's render-phase read was the third instance. Chat + // re-reads the same key post-mount, so the page must not re-grow it. + assert.ok( + !pageSource.includes("readScrollPosition"), + "the page must not read scroll positions at all — the restore lives " + + "in Chat's sessionKey effect (same key, client-only timing)", + ); + }); + + test("the one storage read lives in the post-mount restore effect", () => { + const effectIdx = pageSource.indexOf("const restoredUi = readUiState();"); + assert.ok(effectIdx >= 0, "the mount restore must call readUiState()"); + const effect = pageSource.slice( + pageSource.lastIndexOf("useEffect(", effectIdx), + effectIdx + 400, + ); + assert.match(effect, /readWorkspaceTabs\(\)/, "tabs restore rides the same effect"); + assert.match(effect, /setPersisted\(restoredUi\)/); + assert.match(effect, /setTabState\(restoredTabs\.tabStrip\)/); + assert.match(effect, /setColumnState\(restoredTabs\.columnLayout\)/); + assert.match(effect, /setPanel\(restoredUi\.panel\)/); + assert.match(effect, /setUiRestored\(true\)/, "the write-back gate must open in the same batch"); + }); + + test("every storage write-back is gated on uiRestored", () => { + // Without the gate, the defaults-seeded first effects would overwrite + // the stored payload BEFORE the restore ran — the red-line-3 data + // loss (panel / tabs / appearance gone after a refresh). + const gateCount = ( + pageSource.match(/if \(!uiRestored\) return;/g) ?? [] + ).length; + assert.ok( + gateCount >= 3, + `expected the write-gate in the panel, tabs and lastSessionId mirrors, found ${gateCount}`, + ); + assert.match( + pageSource, + /useEffect\(\(\) => \{\s*if \(!uiRestored\) return;\s*writeUiState\(/, + "the panel mirror must be gated", + ); + assert.match( + pageSource, + /useEffect\(\(\) => \{\s*if \(!uiRestored\) return;\s*writeWorkspaceTabs\(/, + "the workspace-tabs mirror must be gated", + ); + const lastSessionIdx = pageSource.indexOf("lastSessionId: active,"); + assert.ok(lastSessionIdx >= 0); + const lastSessionEffect = pageSource.slice( + pageSource.lastIndexOf("useEffect(", lastSessionIdx), + lastSessionIdx, + ); + assert.match( + lastSessionEffect, + /if \(!uiRestored\) return;/, + "the lastSessionId mirror must be gated — it can fire before the " + + "restore batch and would drop the stored appearance fields", + ); + }); + + test("the first frame still renders the state=null skeleton uniformly", () => { + // The skeleton is what makes server and client renders identical on + // the first frame; the restore must not have traded it for a + // different first-paint path. + assert.match(pageSource, /if \(!state\) \{/); + assert.match(pageSource, //); + }); + + test("the default-seeded state declarations still exist", () => { + // Belt and braces: the two boxes must start from the shared DEFAULT + // constants (identical on server and client), not from undefined. + assert.match(pageSource, /useState\(DEFAULT_UI_STATE\)/); + assert.match( + pageSource, + /useState\(\s*DEFAULT_WORKSPACE_TABS_STATE,?\s*\)/, + ); + }); +}); + +describe("Chat owns the scroll restore (red line: 刷新后滚动位置还在)", () => { + test("the restore reads the persisted key inside the sessionKey effect", () => { + // Chat's effect re-reads `webui:scroll:v1::` whenever + // the session changes and whenever `initialScrollTop` is absent — + // which is now ALWAYS, since the page passes no such prop. Break this + // line and a refresh lands every conversation back at the top. + assert.match( + chatSource, + /const saved = readPersistedScroll\(sessionKey\);/, + "Chat must re-read the persisted scroll position per session key", + ); + assert.match( + chatSource, + /const best = explicit !== null && explicit > 0 \? explicit : saved;/, + "the saved value must be the fallback when no explicit prop arrives", + ); + }); + + test("the page still persists scroll positions through onScrollPersist", () => { + assert.match( + pageSource, + /onScrollPersist=\{\(top\) => \{/, + "the write half of the scroll contract stays on the page", + ); + assert.match(pageSource, /writeScrollPosition\(sessionId, top\)/); + }); +}); diff --git a/packages/webui/webapp/test/send-confirmation.test.ts b/packages/webui/webapp/test/send-confirmation.test.ts index 122685a1..a7203ae6 100644 --- a/packages/webui/webapp/test/send-confirmation.test.ts +++ b/packages/webui/webapp/test/send-confirmation.test.ts @@ -312,7 +312,7 @@ describe("the composer is wired to the probe, not to the deadline", () => { describe("the draft store carries the kind, not a string to match on", () => { test("reset gives a clean record", () => { resetComposerDraftForTests(); - assert.deepEqual(getComposerDraft(), { + assert.deepEqual(getComposerDraft("s1"), { value: "", error: null, errorKind: null, @@ -323,8 +323,8 @@ describe("the draft store carries the kind, not a string to match on", () => { test("the kind and the outcome are independent fields", () => { resetComposerDraftForTests(); - setComposerDraft({ error: "", errorKind: "unconfirmed", unconfirmed: "accepted" }); - assert.equal(getComposerDraft().errorKind, "unconfirmed"); - assert.equal(getComposerDraft().unconfirmed, "accepted"); + setComposerDraft("s1", { error: "", errorKind: "unconfirmed", unconfirmed: "accepted" }); + assert.equal(getComposerDraft("s1").errorKind, "unconfirmed"); + assert.equal(getComposerDraft("s1").unconfirmed, "accepted"); }); }); diff --git a/packages/webui/webapp/test/slash-routing.test.ts b/packages/webui/webapp/test/slash-routing.test.ts index 920aa049..2d1c1f2c 100644 --- a/packages/webui/webapp/test/slash-routing.test.ts +++ b/packages/webui/webapp/test/slash-routing.test.ts @@ -296,14 +296,15 @@ describe("rejected submissions come back into the composer", () => { // merged patch into it is the difference between "the text came // back" and "the text came back until the next re-render". resetComposerDraftForTests(); - setComposerDraft({ value: "", attachments: [] }); + setComposerDraft("s1", { value: "", attachments: [] }); setComposerDraft( - mergeRestoredDraft(getComposerDraft(), { + "s1", + mergeRestoredDraft(getComposerDraft("s1"), { content: "/stop", attachments: [], }), ); - assert.equal(getComposerDraft().value, "/stop"); + assert.equal(getComposerDraft("s1").value, "/stop"); resetComposerDraftForTests(); }); @@ -361,7 +362,7 @@ describe("rejected submissions come back into the composer", () => { const guardEnd = composerSource.indexOf("}", guardIdx + guard.length); const body = composerSource.slice(guardIdx, guardEnd); assert.ok( - /setComposerDraft\(\s*mergeRestoredDraft\(getComposerDraft\(\),\s*restored\)\s*\)/.test( + /setComposerDraft\(\s*dispatchDraftKey,\s*mergeRestoredDraft\(getComposerDraft\(dispatchDraftKey\),\s*restored\),?\s*\)/.test( body, ), `the guarded body must write mergeRestoredDraft(…) back through ` + diff --git a/release/public-source.json b/release/public-source.json index bb65141a..d571fd89 100644 --- a/release/public-source.json +++ b/release/public-source.json @@ -3447,10 +3447,15 @@ "packages/webui/server/cleanup.js", "packages/webui/server/engine/capabilities.js", "packages/webui/server/engine/errors.js", + "packages/webui/server/engine/host.js", "packages/webui/server/engine/index.js", "packages/webui/server/engine/providers/local-runtime-v2.capabilities.js", "packages/webui/server/engine/providers/local-runtime-v2.js", "packages/webui/server/engine/providers/tui-runtime-adapter.js", + "packages/webui/server/engine/session-export.js", + "packages/webui/server/engine/session-reads.js", + "packages/webui/server/engine/session-tree-reads.js", + "packages/webui/server/engine/usage-reads.js", "packages/webui/server/lib/acp-client.js", "packages/webui/server/lib/agent-team-detect.js", "packages/webui/server/lib/agent-team-status.js", @@ -3587,6 +3592,12 @@ "packages/webui/test/lib/engine-catalogue.test.js", "packages/webui/test/lib/engine-provider-sync.test.js", "packages/webui/test/lib/engine/capabilities.test.js", + "packages/webui/test/lib/engine/capability-snapshot.test.js", + "packages/webui/test/lib/engine/host-facade.test.js", + "packages/webui/test/lib/engine/session-export.test.js", + "packages/webui/test/lib/engine/session-reads.test.js", + "packages/webui/test/lib/engine/session-tree-reads.test.js", + "packages/webui/test/lib/engine/usage-reads.test.js", "packages/webui/test/lib/events-concurrency.test.js", "packages/webui/test/lib/events-hash.test.js", "packages/webui/test/lib/events.test.js", @@ -3665,6 +3676,7 @@ "packages/webui/test/routes/protocol.check.mjs", "packages/webui/test/routes/provider-presets.check.mjs", "packages/webui/test/routes/providers.check.mjs", + "packages/webui/test/routes/session-reads.check.mjs", "packages/webui/test/routes/sessions-search.check.mjs", "packages/webui/test/routes/sessions-switch-workspace-follow.check.mjs", "packages/webui/test/routes/sessions-switch.check.mjs", @@ -3937,6 +3949,7 @@ "packages/webui/webapp/test/message-enter-animation.test.ts", "packages/webui/webapp/test/modals-decision-channels.test.ts", "packages/webui/webapp/test/open-file.test.ts", + "packages/webui/webapp/test/page-hydration.test.ts", "packages/webui/webapp/test/plugins-surface.test.ts", "packages/webui/webapp/test/preview-edit.test.ts", "packages/webui/webapp/test/provider-management.test.ts", diff --git a/scripts/test-tmp-leak.check.mjs b/scripts/test-tmp-leak.check.mjs index 6d27f4bf..171aa12c 100644 --- a/scripts/test-tmp-leak.check.mjs +++ b/scripts/test-tmp-leak.check.mjs @@ -241,6 +241,7 @@ const KNOWN_PREFIXES = [ "mcode-webui-d02-router-", "mcode-webui-d02-sse-", "mcode-webui-d1-merge-", + "mcode-webui-engine-snapshot-", "mcode-webui-libsettings-iso-", "mcode-webui-mock-", "mcode-webui-port-fallback-", @@ -281,6 +282,7 @@ const KNOWN_PREFIXES = [ "webui-events-hash-test-", "webui-events-ro-", "webui-events-test-", + "webui-export-facade-", "webui-export-test-", "webui-first-turn-guard-", "webui-lan-gate-test-events-", @@ -316,6 +318,7 @@ const KNOWN_PREFIXES = [ "webui-switch-test-db-", "webui-switch-test-events-", "webui-transcript-test-", + "webui-tree-facade-", "webui-ws-browse-", "webui-ws-gate-", "webui-ws-max-",