diff --git a/docs/webui.md b/docs/webui.md index 30867c91d..14c64c0d6 100644 --- a/docs/webui.md +++ b/docs/webui.md @@ -258,6 +258,33 @@ Two consequences of that table are deliberate rather than incidental: **What the user sees.** The permission-mode selector and the model selector are hidden, not disabled and not accompanied by an error message (`webapp/lib/engine-capabilities.ts`, wired in `webapp/components/composer.tsx`). A toast would report a failure for something the user was never able to do, offer nothing to act on, and reappear on every click. The rule is fail-open: the controls are shown until the declaration positively says the engine cannot do it, so a failed or slow `/api/engine-capabilities` request never removes a working control. +### M3-B10: the model and permission writes move behind the facade (no behaviour change) + +`POST /api/set-model` (#58) and `POST /api/permissions` (#59) are the second half of the model endpoint family; B4 moved its read, this batch moves the write. The reasoning leaves the route and lands in `packages/webui/server/engine/model-writes.js`, where it is named, exported and tested on its inputs. + +**Nothing a client can observe changed.** Every status, response field, field order, warning string and engine push — including which push happens first and what it is allowed to say when it fails — is the one these two endpoints produced before. The boundary is: + +| Concern | Home after B10 | +| --- | --- | +| webui id → engine wire value | `resolveEngineModelConfigValue` | +| variant channel vs effort channel, and the push order each implies | `planModelSelectionPush` | +| the `set_config_option` calls | `pushEngineModelSelection` / `pushEnginePermissionMode` | +| permission mode → label / engine value | `resolvePermissionSelection` | +| the rule for when the local `configOptions` snapshot may claim the engine's new effort | `applyThinkingEffortMirror` (the write stays in the route — `cs` is webui's own state) | +| body parsing, the 400s, the `cs.model` / `cs.permissions` writes, `pushStateFor`, the response bodies | `packages/webui/server/routes/model.js` | + +Two forms the picker deals with are deliberately different and stay that way. What the **engine** receives is the wire form — `m:::u`, or `m:::v:` for a switchable builtin, plus a bare level for `thinkingEffort` and an engine vocabulary word for `permissionMode`. What **webui** records is the user-facing form — `cs.model.name` in `/`, `cs.model.thinking`, `cs.permissions` as a label. The map between them is what the suite pins, field by field, over one row per (engine option shape × request shape) in `packages/webui/test/lib/engine/model-writes.test.js`. + +**The variant channel (ticket 36) is unchanged and now covered by name.** A switchable builtin (`thinking_config.mode: switchable`, e.g. MiniMax-M3) has no engine effort vocabulary — the engine rejects every `thinkingEffort` value for it — and advertises it only as the wire pair `v:thinking` / `v:none-thinking`. Such a pick is therefore **one** `model` push carrying both the model and the on/off level, with no second push at all. Every other model keeps the two-push contract: `model` first, then `thinkingEffort`, because the engine rejects an effort set when no model is selected. A cleared level on the variant channel means the engine's **default** variant, not "off" — the normaliser only knows `on` and `off`. + +**The 4-second SSE race window is unchanged**, and now has both halves pinned. A pick stamps the fields the request actually carried, all with one timestamp, so `applyConfigOptionUpdate`'s ownership-aware mirror (`server/lib/mcode-acp.js`, ticket 08) defers the engine's wire-form echo for 4 seconds instead of letting it overwrite the chip a few milliseconds after the optimistic write. A field the request did *not* carry is not stamped, so a later cross-client change to that field still mirrors immediately. + +**`contextWindow` is still recorded and never pushed.** The engine's ACP surface has no channel for it, so the pick is a webui-side preference the picker reflects immediately. + +**These two endpoints are not gated, and that is an open decision rather than an oversight.** #59 writes `permissionMode` only, so gating it on `authCredentials.setPermissionMode` would be behaviourally inert today and safe against the shipped UI (the permission selector is already hidden under exactly that declaration) — it is one `assertEngineCapability` call. #58 also writes `thinkingEffort`, which is a *generic* config id: gating it the same way would make the thinking-effort control answer 501 for the same reason #68 does for an unrecognised id. Both branches are costed in the KNOWN DEBT section of `model-writes.js` — bridge `thinkingEffort` as a third bridged id, or accept the 501 and extend the frontend's degradation to a third control. Until that is decided, #58 keeps its pre-B10 behaviour. + +**The bridge is no longer an unverified exemption.** `selectModel` and `setPermissionMode` — the two sub-items `MODE_WRITE_BRIDGED_CONFIG_IDS` names — are now in the snapshot audit's `REQUIRED_METHODS`, so a real booted host is checked for both of them on the adapter *and* the CliService surface, and a declaration that stops listing one goes red. Neither surface carries a `setThinkingEffort` / `selectThinkingEffort`, which is the fact the gating decision above turns on. + ### Migration state and constraints - **M1 done in this batch**: host construction (`createCatalogueHost`) moved verbatim into `server/engine/providers/local-runtime-v2.js`; `runtime-host.js` re-exports it, so every existing importer is untouched. No existing route's behaviour changed; `GET /api/engine-capabilities` is a new, additive endpoint. diff --git a/docs/webui.zh-CN.md b/docs/webui.zh-CN.md index 6964dcd7a..1f6af1679 100644 --- a/docs/webui.zh-CN.md +++ b/docs/webui.zh-CN.md @@ -258,6 +258,33 @@ GET /api/engine-capabilities[?provider=] **用户看到什么。** 权限模式选择器与模型选择器被**隐藏**,不是禁用,也不配任何错误提示(`webapp/lib/engine-capabilities.ts`,接线在 `webapp/components/composer.tsx`)。toast 会为一件用户从来就做不到的事报一次失败、无从处理、而且每点一次就再报一次。这条规则是 fail-open 的:控件会一直显示,直到声明明确说引擎做不到——因此一次失败或超时的 `/api/engine-capabilities` 请求绝不会拿掉一个本来能用的控件。 +### M3-B10:模型与权限写搬进引擎门面(行为零变更) + +`POST /api/set-model`(#58)与 `POST /api/permissions`(#59)是模型端点族的另一半:B4 搬了读,本批搬写。推理逻辑离开路由,落进 `packages/webui/server/engine/model-writes.js`,在那里被命名、导出,并按入参测试。 + +**客户端能观察到的一切都没变。** 每个状态码、每个应答字段、字段顺序、警告文案与引擎推送——包括哪一次推送先发生、失败时它被允许说什么——都与本批之前逐字节一致。边界如下: + +| 关注点 | B10 之后的归属 | +| --- | --- | +| webui id → 引擎 wire 值 | `resolveEngineModelConfigValue` | +| variant 通道 vs 强度通道,以及各自蕴含的推送顺序 | `planModelSelectionPush` | +| `set_config_option` 调用 | `pushEngineModelSelection` / `pushEnginePermissionMode` | +| 权限模式 → 标签 / 引擎值 | `resolvePermissionSelection` | +| 本地 `configOptions` 快照何时可以宣称引擎的新强度 | `applyThinkingEffortMirror`(写仍留在路由——`cs` 是 webui 自己的状态) | +| 请求体解析、400、`cs.model` / `cs.permissions` 写入、`pushStateFor`、应答体 | `packages/webui/server/routes/model.js` | + +选择器面对的两种形态本就不同,并且刻意保持不同。**引擎**收到的是 wire 形态——`m:::u`,可切换内置模型则是 `m:::v:`,另加 `thinkingEffort` 的裸档位与 `permissionMode` 的引擎词汇。**webui** 记录的是面向用户的形态——`/` 的 `cs.model.name`、`cs.model.thinking`、作为标签的 `cs.permissions`。两者之间的映射由测试逐字段钉住:`packages/webui/test/lib/engine/model-writes.test.js` 按(引擎选项形态 × 请求形态)每种组合一行。 + +**variant 通道(ticket 36)语义不变,并且现在被具名覆盖。** 可切换内置模型(`thinking_config.mode: switchable`,如 MiniMax-M3)没有引擎强度词汇——引擎会拒绝它的一切 `thinkingEffort` 取值——只以 `v:thinking` / `v:none-thinking` 这一对 wire 形态公布。因此这样的选择是**一次** `model` 推送,同时带上模型与开/关档位,压根没有第二次推送。其余模型保持双推送契约:先 `model` 后 `thinkingEffort`,因为引擎在未选模型时会拒绝设置强度。在 variant 通道上「清空档位」意味着引擎的**默认** variant,而不是「off」——归一化器只认 `on` 与 `off`。 + +**4 秒 SSE 竞态窗口不变**,且两个半边都被钉住。一次选择会为请求**实际携带**的字段打戳,全部共用同一个时间戳,于是 `applyConfigOptionUpdate` 的归属感知镜像(`server/lib/mcode-acp.js`,ticket 08)会把引擎的 wire 形态回声推迟 4 秒,而不是让它在乐观写入后几毫秒就覆盖芯片。请求**未**携带的字段不会被打戳,因此该字段后续的跨端变化仍会立即镜像。 + +**`contextWindow` 依旧只记录、不推送。** 引擎 ACP 面没有它的通道,因此这项选择是 webui 侧的偏好,选择器立刻就能反映。 + +**这两个端点没有挂门,而这是一个待人拍板的开口,不是疏漏。** #59 只写 `permissionMode`,所以把它挂到 `authCredentials.setPermissionMode` 上,今天在行为上是空转的,而且对已发布 UI 安全(权限选择器本来就按同一条声明被隐藏)——那只是一次 `assertEngineCapability` 调用。#58 还会写 `thinkingEffort`,而它是**通用** config id:照样挂门会让思考强度控件因为与 #68 遇到无法识别的 id 时完全相同的原因开始答 501。两个分支的成本都写在 `model-writes.js` 的 KNOWN DEBT 段——把 `thinkingEffort` 桥接成第三个 id,还是接受 501 并把前端降级扩到第三个控件。在拍板之前,#58 保持 B10 之前的行为。 + +**桥接不再是未经核实的豁免。** `selectModel` 与 `setPermissionMode`——`MODE_WRITE_BRIDGED_CONFIG_IDS` 点名的两个子项——现已进入快照审计的 `REQUIRED_METHODS`,因此真实启动的 host 会在 adapter **与** CliService 两个面上被检查这两个方法,而停止列出其中之一的声明会变红。两个面都没有 `setThinkingEffort` / `selectThinkingEffort`,这正是上面那个挂门决策所依据的事实。 + ### 迁移状态与边界 - **本批只做迁移第一步 M1**:host 构造(`createCatalogueHost`)原样移入 `engine/providers/local-runtime-v2.js`,`runtime-host.js` 转发导出,既有引用方零改动;没有任何现有路由行为变化,`GET /api/engine-capabilities` 是纯新增端点。 diff --git a/packages/local-runtime/src/persistence/migration/agent-name-conflict-migration.ts b/packages/local-runtime/src/persistence/migration/agent-name-conflict-migration.ts index 7c0c73e29..734860a42 100644 --- a/packages/local-runtime/src/persistence/migration/agent-name-conflict-migration.ts +++ b/packages/local-runtime/src/persistence/migration/agent-name-conflict-migration.ts @@ -36,10 +36,37 @@ import { const MANIFEST_FILE = 'agent-name-conflicts.json'; const EMPTY_REFERENCE_COUNTS = EMPTY_AGENT_NAME_CONFLICT_REFERENCE_COUNTS; -// Synchronous backup/rewrite needs a long-lived dataDir lease. -export const AGENT_NAME_CONFLICT_MIGRATION_LOCK_STALE_MS = 30 * 60_000; -const AGENT_NAME_CONFLICT_MIGRATION_LOCK_RETRIES = { - retries: 120, +// Synchronous backup/rewrite needs a dataDir lease, so the lease outlives the +// critical section rather than the other way round. +// +// The window is 2 minutes, and it was 30. The reason is not "2 minutes is +// enough to rewrite a database" — a live holder refreshes the lease every +// `stale / 2` (proper-lockfile derives the heartbeat from the stale window), so +// the ratio that actually matters — "two missed heartbeats before the lease is +// called abandoned" — is unchanged. What changed is the cost of a lease that +// is abandoned for real. +// +// This lease gates PROCESS STARTUP: `mcode acp` takes it during the V2 cutover +// and cannot serve a prompt without it. The lease is a bare directory, so a +// process killed between `mkdir` and `release` leaves it behind, and the only +// evidence of a live holder is the directory's mtime. A 30-minute window +// therefore made one killed process unstartable for half an hour — and the +// retry budget (55s) was far too short to outlast it, so every launch during +// that window burned 55s of backoff and then died with +// `agent_name_conflict_migration_failed:lock`. That is the outage: repeated +// sends producing no reply, no engine process, and no session. +// +// The migration itself is idempotent and re-inspects under the lease +// (`ensureAgentNameConflictMigrationLocked`), so the worst a wrongly-considered +// stale lease can cost is one extra inspection — not a double rewrite. +export const AGENT_NAME_CONFLICT_MIGRATION_LOCK_STALE_MS = 2 * 60_000; +// The wait must be able to OUTLAST the stale window, or a waiter gives up at +// the moment the abandoned lease becomes reapable. At this backoff shape the +// 120 retries summed to ~55s against a 30-minute window — a waiter could never +// win, only die. 400 retries sum to ~195s, which rides out one full stale +// expiry and then acquires. +export const AGENT_NAME_CONFLICT_MIGRATION_LOCK_RETRIES = { + retries: 400, factor: 1.2, minTimeout: 25, maxTimeout: 500, diff --git a/packages/local-runtime/test/unit/agent-name-conflict-migration-lock.test.ts b/packages/local-runtime/test/unit/agent-name-conflict-migration-lock.test.ts new file mode 100644 index 000000000..ce02ec463 --- /dev/null +++ b/packages/local-runtime/test/unit/agent-name-conflict-migration-lock.test.ts @@ -0,0 +1,121 @@ +import { mkdir, mkdtemp, rm, utimes } from 'node:fs/promises'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { afterEach, describe, expect, it } from 'vitest'; +import lockfile from 'proper-lockfile'; +import { + AGENT_NAME_CONFLICT_MIGRATION_LOCK_RETRIES, + AGENT_NAME_CONFLICT_MIGRATION_LOCK_STALE_MS, + withAgentNameConflictMigrationLock, +} from '../../src/persistence/migration/agent-name-conflict-migration.js'; + +// The lease that gates `mcode acp` startup. `proper-lockfile` represents it as +// a bare directory (`.lock`): mkdir acquires, rmdir releases, and a +// live holder heartbeats the directory mtime every `stale / 2`. A process +// killed between the two leaves the directory behind, and an abandoned lease is +// indistinguishable from a live one except by that mtime. +// +// These tests pin the two properties the outage depended on. An abandoned lease +// must become reapable on a timescale a process launch can wait out, AND a +// lease that really is held must still be respected. Both halves matter: a fix +// that only shortened the window would trade a startup outage for two +// processes inside one critical section. + +const cleanups: Array<() => Promise> = []; + +afterEach(async () => { + while (cleanups.length > 0) await cleanups.pop()!(); +}); + +async function makeDataDir(): Promise { + const dir = await mkdtemp(join(tmpdir(), 'agent-name-conflict-lock-')); + cleanups.push(async () => { + await rm(dir, { recursive: true, force: true }); + }); + cleanups.push(async () => { + await rm(`${dir}.lock`, { recursive: true, force: true }); + }); + return dir; +} + +/** Create the lease directory and age its mtime to `ageMs` in the past. */ +async function plantAgedLease(dataDir: string, ageMs: number): Promise { + const leaseDir = `${dataDir}.lock`; + await mkdir(leaseDir, { recursive: true }); + const when = new Date(Date.now() - ageMs); + await utimes(leaseDir, when, when); +} + +/** + * An abandoned lease of a fixed, real-world age. + * + * The age is ABSOLUTE and deliberately not derived from the stale constant. + * A fixture aged by `STALE + margin` is reaped instantly under any value of + * the constant, so it passes whatever the constant is — it cannot fail, and a + * test that cannot fail is decoration. Five minutes is the shape of the outage + * this pins: a lease orphaned 15+ minutes earlier by a killed process, still + * blocking every engine launch. + */ +const ABANDONED_LEASE_AGE_MS = 5 * 60_000; + +describe('agent name conflict migration lease', () => { + it('takes over an abandoned lease promptly instead of burning the whole retry budget', async () => { + const dataDir = await makeDataDir(); + await plantAgedLease(dataDir, ABANDONED_LEASE_AGE_MS); + + const startedAt = Date.now(); + await expect(withAgentNameConflictMigrationLock(dataDir, () => 'reaped')).resolves.toBe( + 'reaped', + ); + + // Fast, not merely eventually correct. Before the fix a 30-minute stale + // window made this wait out the whole ~55s retry budget and then throw, + // because the budget could not outlast the window — which is what killed + // every engine launch. + expect(Date.now() - startedAt).toBeLessThan(10_000); + }); + + it('still leaves a lease that is genuinely held alone', async () => { + const dataDir = await makeDataDir(); + // A real holder, taken with the same library the migration uses, so the + // lease directory is a genuine one and its mtime is a live heartbeat — + // which is the ONLY thing that separates a held lease from an abandoned + // one. This is the half that guards against over-correcting: if the stale + // window were shortened to the point of reaping a beating lease, the + // takeover would trade a startup outage for two processes inside one + // critical section. + const release = await lockfile.lock(dataDir, { + stale: AGENT_NAME_CONFLICT_MIGRATION_LOCK_STALE_MS, + }); + + let entered = false; + const acquisition = withAgentNameConflictMigrationLock(dataDir, () => { + entered = true; + }); + // Long enough to prove it is waiting rather than barging in. + await new Promise((resolve) => setTimeout(resolve, 3_000)); + expect(entered).toBe(false); + + // Once the holder releases, the waiter proceeds on its own. + await release(); + await acquisition; + expect(entered).toBe(true); + }); + + it('keeps the retry budget long enough to outlast the stale window', async () => { + // The invariant behind both tests: a waiter must be able to survive one + // abandoned-lease expiry and then acquire, rather than dying at the moment + // the lease becomes reapable. With 120 retries at this backoff shape the + // wait was ~55s against a 30-minute window, so it could only ever lose. + const { retries, factor, minTimeout, maxTimeout } = + AGENT_NAME_CONFLICT_MIGRATION_LOCK_RETRIES; + let total = 0; + let delay = minTimeout; + for (let attempt = 0; attempt < retries; attempt += 1) { + total += delay; + delay = Math.min(delay * factor, maxTimeout); + } + + expect(total).toBeGreaterThan(AGENT_NAME_CONFLICT_MIGRATION_LOCK_STALE_MS); + }); +}); diff --git a/packages/webui/docs/ARCHITECTURE.md b/packages/webui/docs/ARCHITECTURE.md index eba5e21f6..a2f1eba01 100644 --- a/packages/webui/docs/ARCHITECTURE.md +++ b/packages/webui/docs/ARCHITECTURE.md @@ -1778,6 +1778,194 @@ the facade at all. second place that knows provider ids before one exists is precisely the thing M4 exists to remove. This batch deliberately does not create a premature registry. +#### Which endpoints route through the facade (step M3, batch B10) + +Batch B10 takes the write half of the model / permission family — the half B4 left +behind when it moved the read side into `engine/model-reads.js` — into +`engine/model-writes.js`. It is the first write family in this migration with +**zero observable change**: every status, every field and field order, every +warning string and every push order #58 and #59 produce is the one they produced +before the batch, and the suite pins each of them as a value. What moved is +*where the reasoning lives*. The webui-id → engine-wire translation, the +variant-versus-effort channel decision, the two `set_config_option` pushes and +the permission label mapping are now named, exported and testable on their own +inputs instead of being inline branches in a route; `routes/model.js` is net +−100 lines as a result. + +| Endpoint | Facade function | Capability · sub-item | Enforcement | Value source | +| --- | --- | --- | --- | --- | +| `POST /api/set-model` (#58) | `engine/model-writes.js#pushEngineModelSelection` | none declared | **not gated** | `mcode-rpc.js#setConfigOption` at most twice; everything recorded lands in webui's own `cs.model` | +| `POST /api/permissions` (#59) | `engine/model-writes.js#pushEnginePermissionMode` | none declared | **not gated** | `mcode-rpc.js#setConfigOption` once; the recorded label is webui's own `cs.permissions` | + +The concern split inside those two rows is the shape B9 drew for the mode-write +family: the engine-facing half moved, the client-state half stayed. + +| Concern | Home after B10 | +| --- | --- | +| webui model id → the engine's wire value | `engine/model-writes.js#resolveEngineModelConfigValue` | +| the model a request is aimed at | `engine/model-writes.js#modelSelectionTarget` | +| variant channel vs effort channel, and what each push carries | `engine/model-writes.js#planModelSelectionPush` | +| the `set_config_option` pushes, in the plan's order | `engine/model-writes.js#pushEngineModelSelection` | +| permission mode → label **and** engine value | `engine/model-writes.js#resolvePermissionSelection` | +| the permission-mode push | `engine/model-writes.js#pushEnginePermissionMode` | +| the `configOptions` snapshot mirror rule | `engine/model-writes.js#applyThinkingEffortMirror` (the rule here, the write in the route) | +| the `*PickedAt` race stamps | `engine/model-writes.js#planModelPickStamps` | +| body parsing, the 400s, the 200, `cs.model` / `cs.permissions`, `state-bus.js#pushStateFor` | `routes/model.js#handleSetModel` and `routes/model.js#handleSetPermissions` | + +**The id translation exists because the two sides spell a model differently.** +webui records `cs.model.name` in `/` form; the +engine's `model` config id accepts only its own wire encoding, and rejects +anything else. Without the translation a mid-session pick of a multi-segment id +(`nousresearch/deepseek/x`) would 400 from the engine. +`engine/model-writes.js#resolveEngineModelConfigValue` is the seam, and it +returns `null` — rather than guessing — when the engine has no `model` option in +the snapshot yet, which is the state before the first session event lands. The +caller then falls back to the recorded id and `mcode-acp.js#applyRecordedModel` +re-applies it on the next boot, so the mid-session push and the boot-time replay +share one resolver instead of two. + +**The two channels are mutually exclusive, and the order is the engine's +contract.** `engine/model-writes.js#planModelSelectionPush` returns a plan — data, +not a side effect — and the plan has one of two shapes: + +| Channel | When | Pushes | Why | +| --- | --- | --- | --- | +| `variant` | the target is a switchable builtin (the engine advertises `thinking_config.mode: switchable` plus a variant tree) | **one** `model` push carrying both the model and the on/off level; `thinkingPush` is null | such a model has no effort vocabulary at all — the engine rejects every `thinkingEffort` value for it and advertises the level only as part of the model wire value, so a second push has nothing to say | +| `effort` | everything else | a `model` push when the request names a model, then a `thinkingEffort` push when it names a non-empty level | the engine rejects a `thinkingEffort` set while no model is selected, so model first, effort second — the order is a contract, not a style | + +`engine/model-writes.js#modelSelectionTarget` is what makes the effort channel's +model-only request possible: the fallback to the already-recorded model is why a +thinking-only update on a switchable builtin lands at all, and it is exported +rather than inlined so the executor and the planner cannot derive it twice and +drift. + +**What counts as "carried" differs per channel, on purpose.** A plan field, +`carriedThinking`, answers "did *this* push carry a level". On the variant +channel an unchanged recorded level is still carried by the model push, so an +absent `thinking` field falls back to the recorded value. On the effort channel +a level is carried only when the request carried one — an absent field means +"leave the recorded effort alone", and there is no wire form here that could +carry it without also re-selecting the model. A **cleared** field is carried on +neither channel. Collapsing the three into one predicate reads like a +simplification and changes `thinkingSynced` on real, successful pushes, so the +test pins them separately. + +**`mcodeSynced` reports the model, and only the model.** It is false for a +thinking-only update even when that update succeeded, because the field means +"the model is in the engine" and there was no model in the request; +`thinkingSynced` reports the level. On the effort channel a second failure only +escalates the warning when the model push left it untouched, so a model +rejection is never overwritten by the effort rejection it caused — and that is +why a three-way disjunction in the old route collapsed to a two-way one here. + +**The permission endpoint needs two forms of one mode, and the seam that +produces both is the point.** `engine/model-writes.js#resolvePermissionSelection` +answers a label *and* an engine value from one input, because the endpoint needs +both and a fifth form added to one mapper and forgotten in the other is the +failure this prevents. + +| webui mode | label recorded and pushed to every tab (`server/lib/interaction/permission-presets.js#webuiModeToLabel`) | engine value (`mcode-rpc.js#webuiPermissionToMcode`) | +| --- | --- | --- | +| `ask` | Ask | `default` | +| `auto` | Auto | `auto` | +| `read` | Read | `read` | +| `off` | Off | `off` | +| `full` | Full access | `bypassPermissions` | +| anything else | Full access | **null** | + +The last row is load-bearing, not an oversight. The two mappers **disagree on +purpose** about an unrecognised mode: the label mapper falls back to `full` so +the UI always has something to render, while the engine mapper returns null +because there is no engine word for a mode the user invented. So +`POST /api/permissions {"mode":"nonsense"}` records "Full access", pushes +nothing, and answers `mcodeSynced:false` with no warning — and that guard is the +difference between "the engine is in this mode" and "we hope it is". + +**The 4-second window is a two-sided contract, and this batch owns the write +side of it.** The engine's `config_option_update` re-asserts its own wire-form +`currentValue`; without a marker it lands that wire form on the user's pick a +few milliseconds after the optimistic write, and the composer chip flickers +between the friendly recorded form and the engine's. `mcode-acp.js` reads +`modelPickedAt` / `thinkingPickedAt` and defers its mirror while the stamp is +fresh (`mcode-acp.js#PICK_DEFER_WINDOW_MS`, 4000). The reader is not this +batch's to change; `engine/model-writes.js#planModelPickStamps` is the writer's +half, and it carries two properties the suite pins separately: + +| Property | Form | The half of the race it closes | +| --- | --- | --- | +| **one** timestamp for every field of one request | the caller passes `pickAt` in, taken once before the engine is called, and all stamped fields share it by construction | the forward half — a pick that takes 30 ms must not leave the model field expiring 30 ms before the effort field | +| **only** the fields the body actually carried | `modelPickedAt` only when a model was named, `thinkingPickedAt` only when `thinking` was present in the body, `contextWindowPickedAt` only when `contextWindow` was | the reverse half — a thinking-only update must not refresh `modelPickedAt`, or a later cross-client model change is suppressed by a pick the user never made; a "stamp everything" simplification breaks this silently | + +`contextWindowPickedAt` rides along for symmetry with the two fields the mirror +reads. It is recorded and nothing consumes it today, because the engine has no +context-window channel; it was stamped before this batch and stays stamped. + +**The mirror rule is half in the facade and half in the route, and the split is +B9's.** `engine/model-writes.js#applyThinkingEffortMirror` owns the *rule* — +after an accepted effort push the local snapshot should claim the engine's new +value; after a cleared pick it should claim none — and returns how many options +it touched, which is what makes "no `thinkingEffort` option in the snapshot yet" +observable instead of a silent no-op. The *write* stays in the route, because +`cs.configOptions` is webui's own view and is mutated in place exactly as before, +on exactly the same conditions. The three arms: + +| Mirror | When | Local `configOptions` | +| --- | --- | --- | +| `{kind: "set", value}` | a non-empty level was pushed and the engine accepted it | claim the engine's new `currentValue` | +| `{kind: "clear"}` | the level was cleared **and** a model also changed | **drop** the local value — the engine picks its own default for the new model, so a mirror left showing the cleared value would be a state the engine never reported | +| `null` | every other case, including a clear on its own | untouched | + +A clear on its own is deliberately **not** mirrored: the next +`config_option_update` applies it, and dropping locally would invent an engine +state. The clear also does not depend on the model push having succeeded, which +is pre-existing behaviour and is preserved as-is rather than tidied up. + +**Three things this batch records as known debt instead of deciding:** + +1. **Neither endpoint is gated, and that is a decision left open for a human.** + B9's gate already exempts exactly the two config ids these endpoints write — + `model` → `selectModel`, `permissionMode` → `setPermissionMode`, the table in + `engine/mode-writes.js#MODE_WRITE_BRIDGED_CONFIG_IDS` — so both sub-items are known + names and neither needs rediscovering. What stops the gate from being armed + here is one more config id, and it is #58's: + + | Branch | Cost | Benefit | + | --- | --- | --- | + | **(a) bridge** `thinkingEffort` as a third id in `MODE_WRITE_BRIDGED_CONFIG_IDS`, pointed at a sub-item meaning "the dedicated thinking-effort writer" | a third name in a table the frontend mirrors, and a third declaration the snapshot audit must then prove exists on both surfaces — today's probe found no `setThinkingEffort` / `selectThinkingEffort` on either, so the name has to be agreed with the engine team first | #58 becomes gateable on the same table as #59, and the two controls stay symmetric | + | **(b) accept** the 501 and degrade the UI | the thinking-effort control disappears for every provider that denies the generic config write — under M4's ACP provider, most of them — and #58 loses a working half to keep an enrichment; `engine-capabilities.ts` would need a third bridged id for the control to follow the same fail-open rule | the capability declaration stops being a lie about a control that still works | + + `thinkingEffort` is a **generic** config id — the one the plan (§3a, row 68) + says has nowhere to be delivered under a provider with no generic write — so + gating #58 the way #59 could be gated makes the thinking-effort control answer + 501 for the same reason #68 does for an unrecognised id. Until a human picks + a branch, #58 keeps its pre-B10 behaviour. **#59 alone is the zero-risk half:** + gating it hard on `engine/capabilities.js#assertEngineCapability` is behaviourally + inert today (no registered provider lists that sub-item as missing, and the + snapshot audit proves both providers really have the method) and is safe + against the shipped UI, which already hides the permission selector under + exactly that declaration (`webapp/lib/engine-capabilities.ts` + + `webapp/components/composer.tsx`). It is still not taken here, because taking + it would be making a product decision by capability table, with no changelog + and no frontend work — the same argument B7 recorded for #71. The module is + gate-ready either way: the push is one call site per endpoint, so arming + either gate is one line. +2. **`contextWindow` is recorded and never pushed.** The engine's ACP surface has + no channel for it — `session/set_config_option` accepts exactly three config + ids and the model wire encoding has no context segment — so the pick is a + webui-side preference the picker reflects immediately. Pre-existing, unchanged + here, and listed because this batch is the one that owns the whole #58 write: + a reader of the facade should not assume the whole request reaches the + engine. Wiring it is engine-side work, and the seam is the output of + `engine/model-writes.js#planModelSelectionPush`, which a future engine + channel would extend with a third push. +3. **B9's bridge-naming debt is closed by this batch, and the record is the + snapshot audit.** `selectModel` and `setPermissionMode` are now in + `REQUIRED_METHODS`, so + `test/lib/engine/capability-snapshot.test.js#auditProviderCapabilities` + asserts they are functions on both the adapter and the cliService surface of + a real booted host. They were verified present before being added. + `engine/mode-writes.js` is a read-only reference in this batch, so its own + debt text is left exactly as written; this entry is the closure record. ## 6. Frontend topology diff --git a/packages/webui/docs/ARCHITECTURE.zh-CN.md b/packages/webui/docs/ARCHITECTURE.zh-CN.md index 870732174..1779293af 100644 --- a/packages/webui/docs/ARCHITECTURE.zh-CN.md +++ b/packages/webui/docs/ARCHITECTURE.zh-CN.md @@ -1434,6 +1434,161 @@ getter、逐回合 host 包装器与附件辅助函数只通过 `await import()` `MCODE_WEBUI_TRANSPORT` 与字面量 `"runtime"` 比较,而计划书写的是选择应当读 provider 注册表。注册表归 M4 所有,而在它存在之前就硬写第二处知道 provider id 的地方,正是 M4 要消灭的东西。本批刻意不去造一个提前到来的注册表。 +#### 哪些端点经由门面路由(迁移步 M3 批次 B10) + +批次 B10 把模型 / 权限族的写侧收进 `engine/model-writes.js`——正是 B4 把读侧搬进 +`engine/model-reads.js` 时留下的另一半。它是本次迁移里第一个**可观察行为零变化**的 +写族:#58 与 #59 产出的每一个状态码、每一个字段及其顺序、每一条 warning 字符串、 +每一次推送顺序都与本批之前完全相同,测试套件把它们逐个作为取值钉住。变的是 +**推理放在哪里**:webui id → 引擎 wire 值的翻译、variant 与 effort 两条通道的判定、 +两次 `set_config_option` 推送、权限标签映射,如今都是有名、有导出、可单独针对入参 +测试的函数,而不再是路由里的行内分支;`routes/model.js` 因此净减 100 行。 + +| 端点 | 门面函数 | 能力 · 子项 | 强制方式 | 取值来源 | +| --- | --- | --- | --- | --- | +| `POST /api/set-model`(#58) | `engine/model-writes.js#pushEngineModelSelection` | 未声明 | **未挂门** | 至多两次 `mcode-rpc.js#setConfigOption`;被记录的一切都落在 webui 自己的 `cs.model` 里 | +| `POST /api/permissions`(#59) | `engine/model-writes.js#pushEnginePermissionMode` | 未声明 | **未挂门** | 一次 `mcode-rpc.js#setConfigOption`;被记录的标签是 webui 自己的 `cs.permissions` | + +这两行内部的关切切分沿用 B9 为模式写族定下的形状:面向引擎的那一半搬走了,客户端 +状态的那一半留了下来。 + +| 关切 | B10 之后的归属 | +| --- | --- | +| webui 模型 id → 引擎 wire 值 | `engine/model-writes.js#resolveEngineModelConfigValue` | +| 一次请求瞄准的是哪个模型 | `engine/model-writes.js#modelSelectionTarget` | +| variant 通道与 effort 通道,以及各自推送什么 | `engine/model-writes.js#planModelSelectionPush` | +| 按计划顺序发出的 `set_config_option` 推送 | `engine/model-writes.js#pushEngineModelSelection` | +| 权限模式 → 标签**与**引擎值 | `engine/model-writes.js#resolvePermissionSelection` | +| 权限模式推送 | `engine/model-writes.js#pushEnginePermissionMode` | +| `configOptions` 快照镜像规则 | `engine/model-writes.js#applyThinkingEffortMirror`(规则在门面,写入在路由) | +| `*PickedAt` 竞态戳 | `engine/model-writes.js#planModelPickStamps` | +| 请求体解析、各个 400、200、`cs.model` / `cs.permissions`、`state-bus.js#pushStateFor` | `routes/model.js#handleSetModel` 与 `routes/model.js#handleSetPermissions` | + +**id 翻译之所以存在,是因为两侧拼写模型的方式不同。** webui 记录的 +`cs.model.name` 是 `/` 形式,而引擎的 `model` 配置 id +只接受它自己的 wire 编码,其余一律拒绝。没有这层翻译,会话中途选中一个多段 id +(`nousresearch/deepseek/x`)会被引擎 400 掉。 +`engine/model-writes.js#resolveEngineModelConfigValue` 是那道缝,而且它在引擎快照里 +还没有 `model` 选项时返回 `null` 而不是猜一个——那正是首个会话事件落地之前的状态; +调用方随后退回已记录的 id,由 `mcode-acp.js#applyRecordedModel` 在下次启动时重新 +套用,于是会话中途的推送与启动时的回放共用同一个解析器,而不是各有一份。 + +**两条通道互斥,而顺序是引擎的契约。** `engine/model-writes.js#planModelSelectionPush` +返回一个计划——是数据,不是副作用——而计划只有两种形状: + +| 通道 | 何时 | 推送 | 原因 | +| --- | --- | --- | --- | +| `variant` | 目标是可切换内置模型(引擎声明 `thinking_config.mode: switchable` 并给出 variant 树) | **一次** `model` 推送,同时携带模型与开关档位;`thinkingPush` 为 null | 这类模型根本没有档位词汇表——引擎对它拒绝任何 `thinkingEffort` 取值,只把档位作为模型 wire 值的一部分对外声明,因此第二次推送无话可说 | +| `effort` | 其余全部情况 | 请求点名模型时推一次 `model`,请求点名非空档位时再推一次 `thinkingEffort` | 引擎在未选中模型时拒绝设置 `thinkingEffort`,所以先模型、后档位——这是契约而非风格 | + +`engine/model-writes.js#modelSelectionTarget` 正是 effort 通道上「只带档位」的请求得以 +成立的原因:退回当前已记录的模型,正是可切换内置模型上的纯档位更新能够落地的原因; +它被导出而不是内联,是为了让执行器与计划器不会各推一份、彼此漂移。 + +**「本次推送是否携带了档位」按通道分别判定,且是刻意的。** 计划里的 +`carriedThinking` 字段回答的是「**这次**推送有没有带档位」。在 variant 通道上, +即便档位与已记录值相同,它仍由那次模型推送携带,所以缺失的 `thinking` 字段会退回 +已记录值;在 effort 通道上,只有请求本身携带了档位才算携带——字段缺失的含义是 +「别动已记录的 effort」,而这里没有任何 wire 形式能在不同时重选模型的前提下把它带 +过去。被**清空**的字段在两条通道上都不算携带。把这三种情形塌缩成一个判定看起来像 +简化,却会在真实成功的推送上改变 `thinkingSynced`,因此测试分别把它们钉住。 + +**`mcodeSynced` 报告的是模型,且只报告模型。** 对一次纯档位更新,即使该更新成功, +它也是 false,因为这个字段的含义是「模型已在引擎里」,而请求里根本没有模型; +`thinkingSynced` 报告档位。在 effort 通道上,第二次失败只有在模型推送没有动过 +warning 时才升级它,因此模型被拒不会被它自己引发的档位被拒覆盖——这也正是旧路由里 +那个三项析取在这里塌缩为两项判断的原因。 + +**权限端点需要同一个模式的两种形态,而同时产出两者的那道缝才是重点。** +`engine/model-writes.js#resolvePermissionSelection` 从一个入参同时给出标签**与**引擎值, +因为端点两者都需要,而「只给一个映射器加上第五种形态、忘了另一个」正是这道缝要防 +的失败。 + +| webui 模式 | 记录并推给每个标签页的标签(`server/lib/interaction/permission-presets.js#webuiModeToLabel`) | 引擎值(`mcode-rpc.js#webuiPermissionToMcode`) | +| --- | --- | --- | +| `ask` | Ask | `default` | +| `auto` | Auto | `auto` | +| `read` | Read | `read` | +| `off` | Off | `off` | +| `full` | Full access | `bypassPermissions` | +| 任何其他值 | Full access | **null** | + +最后一行是承重的,不是疏漏。两个映射器对无法识别的模式**刻意不一致**:标签映射器 +退回 `full`,好让界面总有东西可渲染;引擎映射器返回 null,因为对于用户自己编出来的 +模式,引擎根本没有对应的词。于是 `POST /api/permissions {"mode":"nonsense"}` 记录下 +"Full access"、什么都不推、回 `mcodeSynced:false` 且不带 warning——而这道守卫正是 +「引擎确实处于这个模式」与「我们希望它是」之间的分界。 + +**4 秒窗口是一份双向契约,本批拥有它的写侧。** 引擎的 `config_option_update` 会重新 +声明它自己的 wire 形态 `currentValue`;没有标记的话,它会在乐观写入后几毫秒把这个 +wire 形态盖到用户的选择上,composer 里的芯片于是会在友好的记录形态与引擎形态之间 +闪烁。`mcode-acp.js` 读取 `modelPickedAt` / `thinkingPickedAt`,并在戳还新鲜时推迟 +镜像(`mcode-acp.js#PICK_DEFER_WINDOW_MS`,4000)。读侧不归本批改动; +`engine/model-writes.js#planModelPickStamps` 是写侧的一半,它带着两条被测试分别钉住 +的性质: + +| 性质 | 形态 | 它堵住竞态的哪一半 | +| --- | --- | --- | +| 一次请求的所有字段共用**一个**时间戳 | 调用方把 `pickAt` 传进来,在调用引擎之前取一次,因此所有被戳字段按构造就共享它 | 正向那一半——一次耗时 30 毫秒的选择,绝不能让模型字段比 effort 字段早 30 毫秒过期 | +| 只戳请求体真正携带的字段 | 点名了模型才写 `modelPickedAt`,请求体里有 `thinking` 才写 `thinkingPickedAt`,有 `contextWindow` 才写 `contextWindowPickedAt` | 反向那一半——一次纯档位更新不得刷新 `modelPickedAt`,否则之后来自其他客户端的模型变更会被一次用户从未做出的选择压制掉;「全部都戳」的简化正是悄无声息地破坏这一半 | + +`contextWindowPickedAt` 是为了与镜像读取的两个字段对称而顺带记录的。今天没有任何东西 +消费它,因为引擎没有上下文窗口通道;本批之前它就已被戳上,本批继续戳。 + +**镜像规则一半在门面、一半在路由,切分沿用 B9。** +`engine/model-writes.js#applyThinkingEffortMirror` 拥有*规则*——档位推送被接受之后, +本地快照应当认领引擎的新值;一次清空选择之后,本地快照应当什么都不认领——并返回它 +改动了多少个选项,这正是「快照里还没有 `thinkingEffort` 选项」成为可观察事件、而 +不是一次静默空操作的原因。*写入*留在路由里,因为 `cs.configOptions` 是 webui 自己 +的视图,且是原地改写、条件与此前逐字相同。三个分支: + +| 镜像 | 何时 | 本地 `configOptions` | +| --- | --- | --- | +| `{kind: "set", value}` | 非空档位已推送且引擎接受了 | 认领引擎新的 `currentValue` | +| `{kind: "clear"}` | 档位被清空**且**模型也发生了变化 | **丢弃**本地取值——引擎会为新模型挑自己的默认值,留着一个显示清空值的镜像,等于宣称一个引擎从未上报过的状态 | +| `null` | 其余全部情况,包括单独的清空 | 不动 | + +单独一次清空被刻意**不**镜像:下一次 `config_option_update` 会应用它,而本地丢弃会 +凭空造出一个引擎状态。该清空同样不依赖模型推送是否成功,这是既有行为,此处原样保留 +而不去「收拾干净」。 + +**本批记为已知债而不予决定的三件事:** + +1. **两个端点都没有挂门,而这是一个留给人决定的问题。** B9 的门控已经豁免了这两个 + 端点写入的**恰好那两个** config id——`model` → `selectModel`、`permissionMode` → + `setPermissionMode`,即 `engine/mode-writes.js#MODE_WRITE_BRIDGED_CONFIG_IDS` 里的那张表 + ——因此两个子项都是已知名字,谁也不需要重新发现。挡住在这里挂门的还有**一个** + config id,而它是 #58 的: + + | 分支 | 代价 | 收益 | + | --- | --- | --- | + | **(a) 桥接**:把 `thinkingEffort` 作为第三个 id 写进 `MODE_WRITE_BRIDGED_CONFIG_IDS`,指向一个含义为「专用的思考档位写入方」的子项 | 在一张前端也要镜像的表里多加一个名字,外加一份快照审计从此必须证明存在的第三项声明——今天的探测在两侧都没找到 `setThinkingEffort` / `selectThinkingEffort`,所以这个名字得先与引擎团队商定 | #58 可以与 #59 共用同一张表挂门,两个控件保持对称 | + | **(b) 接受** 501 并降级界面 | 思考档位控件会对每一个拒绝通用配置写入的 provider 消失——在 M4 的 ACP provider 下是大多数——#58 为保住一项能力增强而失去可用的一半;`engine-capabilities.ts` 还需要第三个被桥接的 id,控件才能遵循同样的 fail-open 规则 | 能力声明不再对一个仍然可用的控件撒谎 | + + `thinkingEffort` 是**通用** config id——正是计划书(§3a 第 68 行)说在「没有通用 + 写入的 provider」下无处投递的那一个——因此用 #59 那样的方式给 #58 挂门,会让 + 思考档位控件因为与 #68 对无法识别的 id 完全相同的理由而回 501。在人选定分支 + 之前,#58 保持本批之前的行为。**#59 单独看是零风险的那一半**:按 + `engine/capabilities.js#assertEngineCapability` 硬门控它,在今天是无行为影响的(没有任何 + 已注册 provider 把该子项列为缺失,而快照审计证明两个 provider 确实都有这个方法), + 且对已发布的界面是安全的——界面本来就在同一份声明下隐藏权限选择器 + (`webapp/lib/engine-capabilities.ts` + `webapp/components/composer.tsx`)。这里 + 仍然没有动手,因为动手就等于用一张能力表、既无变更记录也无前端工作地做出一个 + 产品决定——这正是 B7 为 #71 记下的同一条理由。无论如何这个模块已是挂门就绪的: + 每个端点的推送都只有一个调用点,所以打开任何一个门都是一行的事。 +2. **`contextWindow` 只记录、从不推送。** 引擎的 ACP 面上没有它的通道—— + `session/set_config_option` 只接受三个 config id,而模型 wire 编码里没有上下文 + 段——所以这个选择是一个 webui 侧偏好,选择器会立即反映它。这是既有行为,本批 + 未改;之所以列出来,是因为本批是拥有整个 #58 写侧的那一批:读门面的人不应假定 + 整个请求都抵达了引擎。接线是引擎侧的工作,而那道缝就是 + `engine/model-writes.js#planModelSelectionPush` 的输出——将来的引擎通道会以第三次 + 推送扩展它。 +3. **B9 那条「桥接靠编造子项」的债由本批关闭,记录在快照审计里。** + `selectModel` 与 `setPermissionMode` 现在都在 `REQUIRED_METHODS` 里,于是 + `test/lib/engine/capability-snapshot.test.js#auditProviderCapabilities` 会断言它们 + 在真实启动的 host 上、adapter 与 cliService 两个面上都是函数;加入之前先核实过它们 + 确实存在。本批只是只读引用 `engine/mode-writes.js`,因此它自己的债文本原样保留; + 这一条就是那张关闭凭据。 ## 6. 前端拓扑 diff --git a/packages/webui/server/engine/model-writes.js b/packages/webui/server/engine/model-writes.js new file mode 100644 index 000000000..6bf29e38e --- /dev/null +++ b/packages/webui/server/engine/model-writes.js @@ -0,0 +1,554 @@ +// webui/server/engine/model-writes.js +// +// Migration step M3, batch B10: the MODEL / PERMISSION WRITE family — +// +// #58 POST /api/set-model — pick a model, a thinking level, a context window +// #59 POST /api/permissions — change the session's permission mode +// +// B4 moved the READ half of this endpoint family +// (`engine/model-reads.js`); this moves the WRITE half. Nothing about the +// wire changes: every status, field, ordering and warning string these two +// routes produce is the one they produced before this batch, and the suite +// pins them as values. What changed is WHERE the reasoning lives — the +// model-id translation, the variant-channel decision, the thinking-effort +// mirror rule and the permission label mapping are now named, exported and +// tested on their inputs, instead of being inline branches in a route. +// +// The boundary, stated once: +// +// THE ENGINE-FACING HALF MOVED HERE. EVERYTHING ELSE STAYED IN THE ROUTE. +// +// | Concern | Home after B10 | +// | ----------------------------------------- | ------------------------------------------- | +// | webui id → engine wire value | `resolveEngineModelConfigValue` (here) | +// | variant channel vs effort channel | `planModelSelectionPush` (here) | +// | the two `set_config_option` pushes | `pushEngineModelSelection` (here) | +// | permission mode → label / engine value | `resolvePermissionSelection` (here) | +// | the permission-mode engine push | `pushEnginePermissionMode` (here) | +// | body parsing, the 400s, the 200 | `routes/model.js` | +// | `cs.model` / `cs.permissions` writes | `routes/model.js` (B9's rule, reused) | +// | `pushStateFor` and the response body | `routes/model.js` | +// | the `configOptions` snapshot mirror | `applyThinkingEffortMirror` (rule here, write in the route) | +// +// The last row is the deliberate exception to "the route owns client +// state", and it is the same split B9 drew: this module owns the RULE +// ("after an accepted thinkingEffort push, the local snapshot should claim +// the engine's new value; after a cleared pick it should claim none"), the +// route owns the WRITE (`cs.configOptions` is webui's own view, mutated in +// place exactly as before, and only on the same conditions as before). +// +// What this file deliberately does NOT do: +// +// - It does not gate either endpoint. The decision belongs to a human, +// and the KNOWN DEBT section at the bottom costs both branches: the +// push is one call site per endpoint, so arming either gate is one +// line and nothing else in this file moves. +// - It does not build a host. There is no host on this path. +// - It does not own the transport table or the `501` mapping. B9 owns +// those for #67/#68, and this batch does not duplicate them. +// +// Boot-path weight. `routes/model.js` imports this module directly rather +// than through `engine/index.js`, and that is the same call B4 made for +// `model-reads.js`: this module statically imports `lib/engine-catalogue.js` +// (which reaches `js-yaml`), so re-exporting it from the facade index would +// make `engine/index.js` heavier than the rest of the server's one shared +// import site. The server's own boot cost is unchanged — every module +// involved was already on it through this route. `lib/mcode-rpc.js` (and +// with it the ACP client) is reached through `await import()` inside the +// two data-plane functions, so the same rule every other engine family +// follows holds here too. + +import { resolveModelId, variantChannelFor } from "../lib/engine-catalogue.js"; + +/** + * #58's "no session yet" warning — the record-only path. + * + * A string, not a status: the route answers 200 with this in `warning` so + * the composer can show the pick as local-only until a session exists. It + * is exported because it is a WIRE value — a client-visible English + * sentence with a test that compares it character for character — and + * keeping a second hand-typed copy of it next to the only place that + * produces it is how the two drift apart. + * + * @type {string} + */ +export const NO_SESSION_MODEL_WARNING = "no mcode session yet — recorded for the next one"; + +/** + * #59's "no session yet" warning. Different sentence from #58's on + * purpose — the recorded thing differs (a mode vs a model pick) and the + * wire is byte-compared against the pre-B10 route in the suite. + * + * @type {string} + */ +export const NO_SESSION_PERMISSION_WARNING = "no mcode session yet — applies to the next one"; + +/** + * Translate a webui-recorded model id to the engine's wire form. + * + * The webui records `cs.model.name` in `/` + * form (see `engine/model-reads.js#webuiFullModelId`). The engine's + * `set_config_option` for `configId: "model"` rejects anything that + * isn't the wire form `m:::u` (see + * packages/tui/src/acp/control-state.ts#modelConfigValue / agent.ts + * `parseModelConfigValue`). Without this translation a mid-session + * pick of a multi-segment model id (`nousresearch/deepseek/x`) would + * 400 from the engine. + * + * `resolveModelId` (in `lib/engine-catalogue.js`) owns the resolver — + * it is the same code path `applyRecordedModel` uses on session boot, so + * the mid-session push and the boot-time replay share one source of + * truth. Returns `null` when the engine has no matching option yet + * (the engine configOptions list is empty before the first session + * event lands); the caller falls back to the recorded id and the + * next session event re-attempts the apply via `applyRecordedModel`. + * + * `resolveOpts` (ticket 36) passes straight through to + * `resolveModelId` — today only `preferVariant`, used to fold a + * switchable builtin's on/off level into the model selection. + * + * @param {object} cs The client state; only `configOptions` is read. + * @param {string} modelId The recorded webui id, or a variant target. + * @param {{preferVariant?: string}} [resolveOpts] + * @returns {string|null} The engine's `option.value`, or null. + */ +export function resolveEngineModelConfigValue(cs, modelId, resolveOpts) { + if (!modelId || typeof modelId !== "string") return null; + const allOpts = Array.isArray(cs && cs.configOptions) ? cs.configOptions : []; + const modelOption = allOpts.find((o) => o && o.id === "model"); + if (!modelOption) return null; + return resolveModelId(modelId, modelOption, resolveOpts); +} + +/** + * The model a request is aimed at: the one it names, else the one + * already recorded. + * + * The fallback is what makes a thinking-only update on a switchable + * builtin work — the variant rides the model, and a request that carries + * only a level has to be attached to the model the session already has. + * Exported because the executor needs the target BEFORE it can ask + * `variantChannelFor` whether a plan exists, and deriving it twice from + * two places is how the two copies drift. + * + * @param {object} cs Client state; only `model.name` is read. + * @param {string} [modelId] The requested model, "" when absent. + * @returns {string} + */ +export function modelSelectionTarget(cs, modelId) { + return modelId || (cs && cs.model && cs.model.name) || ""; +} + +/** + * The picks a /api/set-model request carries, as a plan the executor can + * run without re-deriving anything. + * + * A plan is data, not a side effect: the branch structure below is the + * part of #58 that is hardest to read in a route (two channels, three + * fields, four interacting flags), and a route cannot test it. It is a + * pure function of its inputs — `cs` is read, never written. + * + * The two channels are the whole of ticket 36, and they are mutually + * exclusive: + * + * VARIANT — the target rides the variant channel, i.e. it is a + * switchable builtin (the engine's `thinking_config.mode: switchable` + * + variant tree, e.g. MiniMax-M3). Such a model has NO engine effort + * vocabulary: the engine rejects every `thinkingEffort` value for it + * ("Thinking effort is not advertised for the selected model") and + * advertises it only as the wire pair `m:...:v:thinking` / + * `m:...:v:none-thinking`. ONE model push therefore carries both the + * model and the on/off level, and there is no second push at all. + * + * EFFORT — everything else: a model push when the request names a + * model, then a `thinkingEffort` push when the request names a + * non-empty level. The ORDER IS THE ENGINE'S CONTRACT: it rejects a + * `thinkingEffort` set when no model is selected + * (`Select a Session model before changing thinking effort.`, + * agent.ts#1003), so model first, then effort. + * + * `carriedThinking` is the pre-B10 route's "was a level actually carried + * by this push" test, and it differs per channel on purpose. On the + * variant channel an UNCHANGED recorded level is still carried by the + * model push, so an absent `thinking` field falls back to the recorded + * value. On the effort channel a level is carried only when the request + * carried one: an absent field means "leave the recorded effort alone", + * and there is no wire form here that could carry it without also + * re-selecting the model. A CLEARED field is carried on neither channel — + * it is neither a level nor an absence, and `variantPlan.level("")` + * resolves it to the engine's default variant. Pinned per channel because + * collapsing them reads like a simplification and changes `thinkingSynced` + * on real, successful pushes. + * + * @param {object} options + * @param {object} options.cs Client state; `configOptions`, `model.name` + * and `model.thinking` are read, nothing is written. + * @param {string} [options.modelId] The requested model, "" when absent. + * @param {boolean} options.thinkingWasProvided "the field was in the body", + * which is NOT the same as "the field is non-empty": an empty + * string is the documented clear sentinel. + * @param {string} [options.thinking] The requested level, "" to clear. + * @param {{variant: object, defaultLevel: string, level: Function}|null} [options.variantPlan] + * From `variantChannelFor`; null on the effort channel. + * @returns {{channel: "variant"|"effort", target: string, + * modelPush: {value: string}|null, thinkingPush: {value: string}|null, + * reportsModelSynced: boolean, carriedThinking: boolean}} + */ +export function planModelSelectionPush(options = {}) { + const { cs, modelId = "", thinkingWasProvided = false, thinking = "", variantPlan = null } = options; + const target = modelSelectionTarget(cs, modelId); + const recordedThinking = (cs && cs.model && cs.model.thinking) || ""; + const reportsModelSynced = Boolean(modelId); + + if (variantPlan) { + const level = variantPlan.level(thinkingWasProvided ? thinking : recordedThinking); + const value = + resolveEngineModelConfigValue(cs, target, { preferVariant: variantPlan.variant[level] }) ?? target; + return { + channel: "variant", + target, + modelPush: { value }, + thinkingPush: null, + reportsModelSynced, + carriedThinking: thinkingWasProvided ? Boolean(thinking) : Boolean(recordedThinking), + }; + } + + return { + channel: "effort", + target, + modelPush: modelId + ? { value: resolveEngineModelConfigValue(cs, modelId) ?? modelId } + : null, + thinkingPush: thinkingWasProvided && thinking ? { value: thinking } : null, + reportsModelSynced, + carriedThinking: thinkingWasProvided ? Boolean(thinking) : false, + }; +} + +/** + * Which fields a /api/set-model pick stamps, and when. + * + * Ticket 08 (the set-model SSE race): the engine's `config_option_update` + * re-asserts its own wire-form `currentValue`, and without a marker it + * would land that wire form on the user's pick a few milliseconds after + * the optimistic write — the chip flickering between the user-friendly + * recorded form and the engine wire form. `server/lib/mcode-acp.js` reads + * `modelPickedAt` / `thinkingPickedAt` and defers the mirror while the + * stamp is FRESH (`PICK_DEFER_WINDOW_MS`, 4s). That file is not this + * batch's to change; this function is the writer's half of the contract + * and the suite pins the reader's half against it. + * + * Two properties are load-bearing and both are preserved verbatim: + * + * 1. ONE timestamp for every field of one request. The window is a race + * window, not three independent ones — a pick that takes 30ms must + * not leave the model field expiring 30ms before the effort field. + * The caller passes `pickAt` in (taken once, before the engine is + * called) so all stamped fields share it by construction. + * 2. ONLY the fields the request actually carried. A thinking-only + * update must not refresh `modelPickedAt`, or a later cross-client + * model change would be suppressed by a pick the user never made — + * that is the reverse half of the race, and the one a + * "stamp everything" simplification silently breaks. + * + * `contextWindowPickedAt` rides along for symmetry with the two fields + * the mirror reads. It is recorded and nothing consumes it today (the + * engine has no context-window channel, see the route's comment on U6); + * it was stamped before this batch and stays stamped. + * + * @param {object} request Which fields the body carried. + * @param {string} [request.modelId] + * @param {boolean} [request.thinkingWasProvided] + * @param {boolean} [request.contextWindowWasProvided] + * @param {number} pickAt The single timestamp for this request. + * @returns {Record} `{}` when nothing was carried, else + * one entry per carried field, all equal to `pickAt`. + */ +export function planModelPickStamps(request = {}, pickAt) { + const stamps = {}; + if (request.modelId) stamps.modelPickedAt = pickAt; + if (request.thinkingWasProvided) stamps.thinkingPickedAt = pickAt; + if (request.contextWindowWasProvided) stamps.contextWindowPickedAt = pickAt; + return stamps; +} + +/** + * Apply the local `configOptions` mirror rule for one /api/set-model + * push. The RULE lives here; the WRITE is the caller's, because + * `cs.configOptions` is webui's own state. + * + * The mirror exists so a follow-up `/api/models` reads the engine's new + * `currentValue` before the SSE flush lands — the same reason + * `applyRecordedModel` writes `cs.configOptions` on the boot path. The + * two arms are the two outcomes: + * + * `{kind: "set", value}` — the engine accepted a new effort; claim it. + * `{kind: "clear"}` — the effort was cleared AND a model changed; + * the engine picks its own default for the new model, so the local + * mirror is DROPPED rather than left showing the cleared value. + * `null` — nothing to do (every other case). + * + * Mutates the array in place and returns how many options it touched, + * which is what makes the "no thinkingEffort option in the snapshot yet" + * case observable instead of a silent no-op. + * + * @param {object[]|undefined} configOptions `cs.configOptions`. + * @param {{kind: "set", value: string}|{kind: "clear"}|null} mirror + * @returns {number} Options changed. + */ +export function applyThinkingEffortMirror(configOptions, mirror) { + if (!mirror) return 0; + const opts = Array.isArray(configOptions) ? configOptions : []; + let touched = 0; + for (const o of opts) { + if (!o || o.id !== "thinkingEffort") continue; + if (mirror.kind === "clear") delete o.currentValue; + else o.currentValue = mirror.value; + touched++; + } + return touched; +} + +/** + * A /api/permissions mode, resolved to both forms the endpoint needs: + * the webui label it records and pushes to every tab, and the engine + * value it forwards. + * + * The five webui ids (`ask` / `auto` / `read` / `off` / `full`, plus any + * unknown or missing one, which both mappers resolve to the `full` + * entry) are mapped in `lib/interaction/permission-presets.js` and + * `lib/mcode-rpc.js` respectively. This function is the single seam that + * says the endpoint needs BOTH, so a future fifth form cannot be added to + * one mapper and forgotten in the other. + * + * Async because one of the two mappers lives behind the RPC wrapper, and + * the RPC wrapper is reached through `await import()` on this module's + * boot-path rule. It is still a function of its input alone. + * + * @param {string} mode The request's `mode`, any case. + * @returns {Promise<{label: string, mcodeValue: string|null}>} + */ +export async function resolvePermissionSelection(mode) { + const [presets, rpc] = await Promise.all([ + import("../lib/interaction/permission-presets.js"), + import("../lib/mcode-rpc.js"), + ]); + const webuiMode = (mode || "full").toLowerCase(); + return { + label: presets.webuiModeToLabel(webuiMode), + mcodeValue: rpc.webuiPermissionToMcode(webuiMode), + }; +} + +/** + * #58 — push the planned model selection to the engine. + * + * The order below IS the endpoint's contract and none of it is new: + * + * 1. NO SESSION → answer with the local-only warning and stop. The pick + * is recorded by the route and re-applied on the next boot by + * `applyRecordedModel`; there is nothing to push and nothing to say + * about `mcodeSynced` beyond false. + * 2. PLAN. `variantChannelFor` reads the engine's materialised builtin + * tree; a plan comes back for a switchable builtin and null for + * everything else. + * 3. PUSH, in the plan's order. The first failure sets the warning; a + * second failure on the effort channel only escalates when the + * warning is still the untouched default, so a model rejection is + * not overwritten by the effort rejection it caused. + * 4. MIRROR DECISION, returned rather than applied (see the module + * header). + * + * `mcodeSynced` reports the MODEL push only, and is false for a + * thinking-only update even when that update succeeded — the field's + * meaning is "the model is in the engine", and there was no model in the + * request. `thinkingSynced` reports the LEVEL. + * + * @param {object} options + * @param {object} options.cs Client state; read only. + * @param {string} [options.cid] Routed to the RPC wrapper, which pins + * the call on the client that owns this tab's session. + * @param {string} [options.modelId] + * @param {boolean} [options.thinkingWasProvided] + * @param {string} [options.thinking] + * @returns {Promise<{channel: string, mcodeSynced: boolean, + * thinkingSynced: boolean, warning: string|null, + * thinkingMirror: {kind: "set", value: string}|{kind: "clear"}|null, + * plan: object}>} + */ +export async function pushEngineModelSelection(options = {}) { + const { cs, cid, modelId = "", thinkingWasProvided = false, thinking = "" } = options; + const sid = cs && cs.mcodeSessionId; + if (!sid) { + return { + channel: "no-session", + mcodeSynced: false, + thinkingSynced: false, + warning: NO_SESSION_MODEL_WARNING, + thinkingMirror: null, + plan: null, + }; + } + const [rpc] = await Promise.all([import("../lib/mcode-rpc.js")]); + const variantPlan = variantChannelFor(modelSelectionTarget(cs, modelId)); + const plan = planModelSelectionPush({ cs, modelId, thinkingWasProvided, thinking, variantPlan }); + + let mcodeSynced = false; + let thinkingSynced = false; + let warning = null; + + if (plan.channel === "variant") { + const r = await rpc.setConfigOption(sid, "model", plan.modelPush.value, cid); + mcodeSynced = plan.reportsModelSynced ? r.ok : false; + thinkingSynced = Boolean(r.ok) && plan.carriedThinking; + if (!r.ok) warning = r.error; + return { channel: "variant", mcodeSynced, thinkingSynced, warning, thinkingMirror: null, plan }; + } + + if (plan.modelPush) { + const r = await rpc.setConfigOption(sid, "model", plan.modelPush.value, cid); + mcodeSynced = r.ok; + if (!r.ok) warning = r.error; + } + if (plan.thinkingPush) { + const r = await rpc.setConfigOption(sid, "thinkingEffort", plan.thinkingPush.value, cid); + thinkingSynced = r.ok; + // Escalate only when the model push left the warning untouched. The + // pre-B10 route spelled this as a three-way disjunction + // (`!warning || warning === null || warning === NO_SESSION_MODEL_WARNING`); + // `warning` is `null` here or a string the model push already set — + // this executor returns the no-session case before reaching the push — + // so `!warning` is the same test without the branch that can never be + // taken. + if (!r.ok && !warning) warning = r.error; + } + // The two mirror arms, and the conditions are the pre-B10 ones: an + // accepted effort claims the engine's new value, while a CLEARED + // effort is mirrored by DROPPING the local value (and only when a + // model also changed — a clear on its own is applied by the next + // `config_option_update`, and dropping here would invent an engine + // state the engine never reported). The clear does NOT depend on the + // model push having succeeded, which is also pre-existing. + const thinkingMirror = + thinkingWasProvided && thinking + ? thinkingSynced + ? { kind: "set", value: thinking } + : null + : thinkingWasProvided && !thinking && modelId + ? { kind: "clear" } + : null; + return { channel: "effort", mcodeSynced, thinkingSynced, warning, thinkingMirror, plan }; +} + +/** + * #59 — push the permission mode to the engine. + * + * One push, one shape. The two conditions that guard it are the + * pre-B10 ones: no session means the change is local until the next one + * (the warning says so), and a mode with no engine value is recorded and + * not pushed. + * + * The second condition is NOT hypothetical. The two mappers disagree on + * an unrecognised mode on purpose: `webuiModeToLabel` falls back to + * `full` so the UI always has a label, while `webuiPermissionToMcode` + * returns null because there is no engine word for a mode the user + * invented. So `POST /api/permissions {"mode":"nonsense"}` records + * "Full access" and pushes nothing — and the guard is the difference + * between "the engine is in this mode" and "we hope it is". + * + * @param {object} options + * @param {object} options.cs Client state; only `mcodeSessionId` is read. + * @param {string} options.mcodeValue From `resolvePermissionSelection`. + * @param {string} [options.cid] + * @returns {Promise<{mcodeSynced: boolean, warning: string|null}>} + */ +export async function pushEnginePermissionMode(options = {}) { + const { cs, cid, mcodeValue } = options; + const sid = cs && cs.mcodeSessionId; + if (!sid) { + return { mcodeSynced: false, warning: NO_SESSION_PERMISSION_WARNING }; + } + if (!mcodeValue) { + return { mcodeSynced: false, warning: null }; + } + const [rpc] = await Promise.all([import("../lib/mcode-rpc.js")]); + const r = await rpc.setConfigOption(sid, "permissionMode", mcodeValue, cid); + return { mcodeSynced: Boolean(r.ok), warning: r.ok ? null : r.error }; +} + +// --------------------------------------------------------------------------- +// KNOWN DEBT +// --------------------------------------------------------------------------- +// +// 1. NEITHER ENDPOINT IS GATED, AND THAT IS A DECISION LEFT OPEN FOR A +// HUMAN — not an oversight. B9's gate already exempts exactly the two +// config ids these endpoints write (`model` → `selectModel`, +// `permissionMode` → `setPermissionMode`, see +// `MODE_WRITE_BRIDGED_CONFIG_IDS` in `engine/mode-writes.js`), so +// both sub-items are known names and neither needs rediscovering. +// What stops the gate from being switched on here is one more config +// id, and it is #58's: +// +// - #59 /api/permissions writes `permissionMode` ONLY. Gating it +// hard on `authCredentials.setPermissionMode` is behaviourally +// inert today (no registered provider lists that sub-item as +// missing, and the snapshot audit now proves both providers +// really have the method) and is safe against the shipped UI, +// which already hides the permission selector under exactly that +// declaration (`webapp/lib/engine-capabilities.ts` + +// `composer.tsx`). The change is one +// `assertEngineCapability(...)` call before the push. +// +// - #58 /api/set-model ALSO writes `thinkingEffort`, and +// `thinkingEffort` is a GENERIC config id — the one the plan +// (§3a, row 68) says has nowhere to be delivered under a +// provider with no generic write. Gating #58 the same way makes +// the thinking-effort control answer 501 for the same reason #68 +// does for an unrecognised id. +// +// Two branches, both costed, neither chosen here: +// +// (a) BRIDGE `thinkingEffort` as a THIRD id in +// `MODE_WRITE_BRIDGED_CONFIG_IDS`, pointed at a sub-item that +// means "the dedicated thinking-effort writer". Cost: a third +// name in a table the frontend mirrors, and a third declaration +// the snapshot audit must then prove exists on both surfaces +// (today's probe found no `setThinkingEffort` / +// `selectThinkingEffort` on either, so the name would have to +// be agreed with the engine team first). Benefit: #58 becomes +// gateable on the same table as #59, and the two controls stay +// symmetric. +// +// (b) ACCEPT the 501 and degrade the UI. Cost: the thinking-effort +// control disappears for any provider that denies the generic +// config write — which, under M4's ACP provider, is most of +// them — and `#58` loses a working half to keep an enrichment. +// `webapp/lib/engine-capabilities.ts` would need a third +// bridged id for the effort control to follow the same +// fail-open rule rather than a 501 at click time. +// +// Until a human picks one, #58 keeps its pre-B10 behaviour, and this +// module stays gate-ready: the push is already a single call site per +// endpoint, so arming either gate is one line in the executor. +// +// 2. B9's KNOWN DEBT 2 (the bridge naming sub-items no audited host was +// proven to have) IS CLOSED BY THIS BATCH, and the evidence is in +// `test/lib/engine/capability-snapshot.test.js`: `selectModel` and +// `setPermissionMode` are now in `REQUIRED_METHODS`, so the audit +// asserts they are functions on BOTH the adapter and the cliService +// surface of a real booted host. They were verified present before +// being added. `engine/mode-writes.js` is a read-only reference in +// this batch, so its own debt text is left as written; this entry is +// the closure record. +// +// 3. `contextWindow` IS RECORDED AND NEVER PUSHED. The engine's ACP +// surface has no channel for it (`session/set_config_option` accepts +// exactly three config ids and the model wire encoding has no context +// segment), so the pick is a webui-side preference the picker +// reflects immediately. That is pre-existing and unchanged here; it +// is listed because this batch is the one that owns the whole +// #58 write, and a reader of this file should not assume the whole +// request reaches the engine. Wiring it is engine-side work; the seam +// is `planModelSelectionPush`'s output, which a future engine +// channel would extend with a third push. diff --git a/packages/webui/server/routes/model.js b/packages/webui/server/routes/model.js index ef1526e02..eb8268f23 100644 --- a/packages/webui/server/routes/model.js +++ b/packages/webui/server/routes/model.js @@ -1,32 +1,38 @@ // webui/server/routes/model.js // GET /api/models, POST /api/set-model, POST /api/permissions, POST /api/answer (legacy) // -// M3-B4: `GET /api/models` now reads the catalogue through the engine -// facade (`server/engine/model-reads.js`) instead of assembling it -// here. The three sources (the engine session's `model` config option, -// the merged providers config with the engine's `custom_provider` -// tree as its bottom layer, the builtin cli-bundle extraction), the -// two builtin-tree annotations (variant-style thinking levels and -// context-window options) and the three derived "what is active" figures -// all moved with it, as named pure functions pinned on their inputs. +// M3-B4 moved the READ (`GET /api/models`) behind the engine facade +// (`server/engine/model-reads.js`). M3-B10 moved the WRITE half: the +// model-id translation, the variant-channel decision, the two +// `set_config_option` pushes and the permission label mapping now live in +// `server/engine/model-writes.js`. // -// The response is byte-identical. This batch only moves the READ: the -// WRITE half (`handleSetModel`) stays here for B7/B9, together with the -// two other handlers below. +// What stayed here, and why: the body parsing and its 400s (caller +// confusion, not an engine limitation), the `cs.model` / `cs.permissions` +// writes, `pushStateFor`, and the response bodies. The mirror rule for +// `cs.configOptions` is the one seam that is half-and-half — the RULE +// (`applyThinkingEffortMirror`) is the facade's, the WRITE stays here, +// because `cs` is webui's own state. See that module's header for the +// full boundary table. +// +// The wire is byte-identical to the pre-B10 route: every status, field +// order, warning string and push order is pinned as a value in +// `test/lib/engine/model-writes.test.js` and, end to end through this +// route, in `test/routes/model.check.mjs`. import { readFileSync } from "node:fs"; import { join } from "node:path"; import { pushStateFor } from "../lib/state-bus.js"; -import { - mcodePermissionToWebui, - setConfigOption, - webuiPermissionToMcode, - PERMISSION_MODES, -} from "../lib/mcode-rpc.js"; +import { mcodePermissionToWebui, PERMISSION_MODES } from "../lib/mcode-rpc.js"; import { readEngineModelCatalogue } from "../engine/model-reads.js"; -import { variantChannelFor, resolveModelId } from "../lib/engine-catalogue.js"; -import { webuiModeToLabel } from "../lib/interaction/permission-presets.js"; +import { + applyThinkingEffortMirror, + planModelPickStamps, + pushEngineModelSelection, + pushEnginePermissionMode, + resolvePermissionSelection, +} from "../engine/model-writes.js"; import { readJson } from "../lib/read-json.js"; /** @@ -58,38 +64,6 @@ function readModelsConfig() { } } -/** - * Translate a webui-recorded model id to the engine's wire form. - * - * The webui records `cs.model.name` in `/` - * form (see `engine/model-reads.js#webuiFullModelId`). The engine's - * `set_config_option` for `configId: "model"` rejects anything that - * isn't the wire form `m:::u` (see - * packages/tui/src/acp/control-state.ts#modelConfigValue / agent.ts - * `parseModelConfigValue`). Without this translation a mid-session - * pick of a multi-segment model id (`nousresearch/deepseek/x`) would - * 400 from the engine. - * - * `resolveModelId` (in `lib/engine-catalogue.js`) owns the resolver — - * it is the same code path `applyRecordedModel` uses on session boot, so - * the mid-session push and the boot-time replay share one source of - * truth. Returns `null` when the engine has no matching option yet - * (the engine configOptions list is empty before the first session - * event lands); the caller falls back to the recorded id and the - * next session event re-attempts the apply via `applyRecordedModel`. - * - * `resolveOpts` (ticket 36) passes straight through to - * `resolveModelId` — today only `preferVariant`, used to fold a - * switchable builtin's on/off level into the model selection. - */ -function translateWebuiModelIdToEngineValue(cs, modelId, resolveOpts) { - if (!modelId || typeof modelId !== "string") return null; - const allOpts = Array.isArray(cs && cs.configOptions) ? cs.configOptions : []; - const modelOption = allOpts.find((o) => o && o.id === "model"); - if (!modelOption) return null; - return resolveModelId(modelId, modelOption, resolveOpts); -} - /** * GET /api/models — the composer model picker, through the engine * facade. @@ -211,92 +185,21 @@ export async function handleSetModel(req, res, ctx) { // would re-assert its wire-form `currentValue` over the user's // recorded pick a few ms after the optimistic write, causing the // chip to flicker between user-friendly form and engine wire form. + // One timestamp for the whole request, and only the fields the body + // actually carried — see `planModelPickStamps`. const pickAt = Date.now(); - if (modelId) cs.model.modelPickedAt = pickAt; - if (thinkingWasProvided) cs.model.thinkingPickedAt = pickAt; - if (contextWindowWasProvided) cs.model.contextWindowPickedAt = pickAt; - const sid = cs.mcodeSessionId; - let mcodeSynced = false; - let thinkingSynced = false; - let warning = sid ? null : "no mcode session yet — recorded for the next one"; - // Ticket 36 — variant channel. Switchable builtin models (the - // engine's `thinking_config.mode: switchable` + variant tree, - // e.g. MiniMax-M3) have NO engine effort vocabulary: the engine - // rejects every `thinkingEffort` value for them ("Thinking effort - // is not advertised for the selected model"). Their on/off level - // rides the MODEL selection instead — the engine advertises such - // models only as variant wire forms (`m:...:v:thinking` / - // `m:...:v:none-thinking`). When the target model rides the - // variant channel, one model push carries both the model and the - // level; the thinkingEffort push below is skipped entirely. - const variantTarget = modelId || (cs.model && cs.model.name) || ""; - const variantPlan = sid ? variantChannelFor(variantTarget) : null; - // Engine contract: model first, then thinkingEffort (the engine - // rejects a thinkingEffort set when no model is selected). Only push - // when BOTH the recorded model and the new (or unchanged) thinking - // are concrete — the engine will validate the level against the - // selected model's effortOptions and reject unknown values. - if (sid && variantPlan) { - const level = variantPlan.level( - thinkingWasProvided ? thinking : cs.model && cs.model.thinking, - ); - const engineValue = - translateWebuiModelIdToEngineValue(cs, variantTarget, { - preferVariant: variantPlan.variant[level], - }) ?? variantTarget; - const r = await setConfigOption(sid, "model", engineValue, ctx.cid); - if (modelId) mcodeSynced = r.ok; - // "thinking synced" reports the level actually carried by the - // push: an explicit pick, or a previously recorded one. An - // engine-default variant (no user-chosen level) is not a sync. - const carriedLevel = thinkingWasProvided ? !!thinking : !!(cs.model && cs.model.thinking); - thinkingSynced = r.ok && carriedLevel; - if (!r.ok) warning = r.error; - } else if (sid) { - if (modelId) { - // The engine wire form is `m:::u` - // (see packages/tui/src/acp/control-state.ts#modelConfigValue). - // The webui id is `/` — translate it - // to the engine wire form so `parseModelConfigValue` accepts it. - // The same resolver used by `applyRecordedModel` lives in - // `lib/mcode-acp.js#resolveModelId` and exports the helper we - // need; the route layer keeps the apply path's reasoning - // (single source of truth for "recorded → engine option.value"). - const engineValue = translateWebuiModelIdToEngineValue(cs, modelId) ?? modelId; - const r = await setConfigOption(sid, "model", engineValue, ctx.cid); - mcodeSynced = r.ok; - if (!r.ok) warning = r.error; - } - if (thinkingWasProvided && thinking) { - const r = await setConfigOption(sid, "thinkingEffort", thinking, ctx.cid); - thinkingSynced = r.ok; - if (!r.ok && (!warning || warning === null || warning === "no mcode session yet — recorded for the next one")) { - warning = r.error; - } - if (r.ok) { - // Mirror the apply on the local configOptions snapshot so a - // follow-up /api/models reads the engine's new currentValue - // before the SSE flush lands (same reason as - // applyRecordedModel's cs.configOptions write). - const opts = Array.isArray(cs.configOptions) ? cs.configOptions : []; - for (const o of opts) { - if (o && o.id === "thinkingEffort") { - o.currentValue = thinking; - } - } - } - } else if (thinkingWasProvided && !thinking && modelId) { - // Model changed AND effort cleared. The engine picks its own - // default for the new model; we drop the local mirror so a - // subsequent /api/models doesn't keep showing the cleared value. - const opts = Array.isArray(cs.configOptions) ? cs.configOptions : []; - for (const o of opts) { - if (o && o.id === "thinkingEffort") { - delete o.currentValue; - } - } - } - } + Object.assign( + cs.model, + planModelPickStamps({ modelId, thinkingWasProvided, contextWindowWasProvided }, pickAt), + ); + const { mcodeSynced, thinkingSynced, warning, thinkingMirror } = await pushEngineModelSelection({ + cs, + cid, + modelId, + thinkingWasProvided, + thinking, + }); + applyThinkingEffortMirror(cs.configOptions, thinkingMirror); pushStateFor(cid); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); return res.end( @@ -315,21 +218,18 @@ export async function handleSetModel(req, res, ctx) { // POST /api/permissions — mid-session permission mode change, through // session/set_config_option{configId:'permissionMode'}. // body: { mode: 'ask'|'auto'|'read'|'full' 或 mcode 原值 } +// +// `resolvePermissionSelection` is the one seam that produces both forms +// of the mode — the webui label recorded on `cs.permissions` and pushed +// to every tab, and the engine value forwarded — so the two mappers +// cannot drift apart. The push itself, and the "no session yet" warning, +// belong to `engine/model-writes.js#pushEnginePermissionMode`. export async function handleSetPermissions(req, res, ctx) { const cs = ctx.cs; const cid = ctx.cid; const payload = await readJson(req); - const webuiMode = (payload.mode || "full").toLowerCase(); - const label = webuiModeToLabel(webuiMode); - const mcodeValue = webuiPermissionToMcode(webuiMode); - const sid = cs.mcodeSessionId; - let mcodeSynced = false; - let warning = sid ? null : "no mcode session yet — applies to the next one"; - if (sid && mcodeValue) { - const r = await setConfigOption(sid, "permissionMode", mcodeValue, ctx.cid); - mcodeSynced = r.ok; - if (!r.ok) warning = r.error; - } + const { label, mcodeValue } = await resolvePermissionSelection(payload.mode); + const { mcodeSynced, warning } = await pushEnginePermissionMode({ cs, cid, mcodeValue }); cs.permissions = label; pushStateFor(cid); res.writeHead(200, { "Content-Type": "application/json; charset=utf-8" }); diff --git a/packages/webui/test/lib/engine/capability-snapshot.test.js b/packages/webui/test/lib/engine/capability-snapshot.test.js index f8412d3b2..16c8cfe1c 100644 --- a/packages/webui/test/lib/engine/capability-snapshot.test.js +++ b/packages/webui/test/lib/engine/capability-snapshot.test.js @@ -101,6 +101,19 @@ function resolveMember(host, dottedPath) { * claim). They are part of the snapshot so "missing must really be * absent" is checked, and a partial that stops listing one goes red * (under-declaration). + * + * M3-B10 added `selectModel` and `setPermissionMode` to `authCredentials` + * on BOTH surfaces. They are the two sub-items + * `MODE_WRITE_BRIDGED_CONFIG_IDS` (server/engine/mode-writes.js) names, + * and until this batch they were the one part of a hard gate that no + * audit could check: `absent` proves a name is NOT on the surface, and a + * name that is merely "not in `missing`" proves nothing. Both were + * verified present by reflection on a booted host BEFORE being added + * here, and the live audit below keeps proving it — which closes the + * bridge question B9 recorded as its KNOWN DEBT 2. Neither surface + * carries a `setThinkingEffort` / `selectThinkingEffort`; that absence + * is the fact the B10 KNOWN DEBT about gating `/api/set-model` turns on, + * and a surface that grows one must add it here at the same time. */ const REQUIRED_METHODS = { "tui-runtime-adapter": { @@ -113,7 +126,7 @@ const REQUIRED_METHODS = { mcp: { on: "adapter", methods: ["configureSessionMcpServers", "clearSessionMcpServers", "inspectProjectMcp", "listMcpServers"] }, subagents: { on: "adapter", methods: ["getDelegationSnapshot", "stopDelegation", "listBackgroundTasks"] }, usageStats: { on: "adapter", methods: ["getSessionUsage", "getSessionUsageSummary", "watchSessionUsageCommits"] }, - authCredentials: { on: "adapter", methods: ["getAccountStatus", "getCodexOAuthStatus", "startCodexOAuthLogin", "cancelCodexOAuthLogin", "getMiniMaxApiKeyStatus", "upsertMiniMaxApiKey", "listUserModelProviders", "createUserModelProvider", "updateUserModelProvider", "deleteUserModelProvider", "testUserModelProvider", "discoverUserModelsCandidate"], absent: ["setConfigOption"] }, + authCredentials: { on: "adapter", methods: ["getAccountStatus", "getCodexOAuthStatus", "startCodexOAuthLogin", "cancelCodexOAuthLogin", "getMiniMaxApiKeyStatus", "upsertMiniMaxApiKey", "selectModel", "setPermissionMode", "listUserModelProviders", "createUserModelProvider", "updateUserModelProvider", "deleteUserModelProvider", "testUserModelProvider", "discoverUserModelsCandidate"], absent: ["setConfigOption"] }, fileReadWrite: { on: "adapter", methods: ["listWorkspaceFileTree", "searchWorkspaceFiles"] }, gitOperations: { on: "adapter", methods: ["getWorkspaceGitMetadata"] }, }, @@ -128,7 +141,7 @@ const REQUIRED_METHODS = { mcp: { on: "cliService", methods: ["configureSessionMcpServers", "inspectProjectMcp", "clearSessionMcpServers", "listMcpServers"] }, subagents: { on: "cliService", methods: ["listBackgroundTasks"], absent: ["getDelegationSnapshot", "stopDelegation"] }, usageStats: { on: "cliService", methods: ["getSessionUsage", "getSessionUsageSummary", "watchSessionUsageCommits"] }, - authCredentials: { on: "cliService", methods: ["getAccountStatus", "getCodexOAuthStatus", "startCodexOAuthLogin", "cancelCodexOAuthLogin", "getMiniMaxApiKeyStatus", "upsertMiniMaxApiKey", "listUserModelProviders", "createUserModelProvider", "updateUserModelProvider", "deleteUserModelProvider", "testUserModel", "discoverUserModelsCandidate"], absent: ["setConfigOption"] }, + authCredentials: { on: "cliService", methods: ["getAccountStatus", "getCodexOAuthStatus", "startCodexOAuthLogin", "cancelCodexOAuthLogin", "getMiniMaxApiKeyStatus", "upsertMiniMaxApiKey", "selectModel", "setPermissionMode", "listUserModelProviders", "createUserModelProvider", "updateUserModelProvider", "deleteUserModelProvider", "testUserModel", "discoverUserModelsCandidate"], absent: ["setConfigOption"] }, fileReadWrite: { on: "cliService", methods: ["listWorkspaceFileTree", "searchWorkspaceFiles"] }, gitOperations: { on: "cliService", methods: ["getWorkspaceGitMetadata", "getWorkspaceReviewLink"] }, }, diff --git a/packages/webui/test/lib/engine/model-writes.test.js b/packages/webui/test/lib/engine/model-writes.test.js new file mode 100644 index 000000000..5eb3879d7 --- /dev/null +++ b/packages/webui/test/lib/engine/model-writes.test.js @@ -0,0 +1,996 @@ +// webui/test/lib/engine/model-writes.test.js +// +// M3-B10 — the MODEL / PERMISSION WRITE family (#58 set-model, +// #59 permissions). +// +// B10 is a MOVE, so this file is organised around the equivalence claim +// rather than around the code: every case names which side of it it +// pins. +// +// THE WIRE FORM. What the ENGINE receives — +// `m:::u` or `m:::v:` for +// the model, a bare level for `thinkingEffort`, an engine vocabulary +// word for `permissionMode`. Pushed through the same +// `lib/mcode-rpc.js#setConfigOption` wrapper, in the same order, with +// the same warnings, as the pre-B10 route. +// +// THE RECORDED FORM. What WEBUI keeps — `cs.model.name` in +// `/`, `cs.model.thinking`, the +// `*PickedAt` stamps, `cs.permissions` as a label — plus the +// `mcodeSynced` / `thinkingSynced` / `warning` triple the response +// body carries. +// +// The two are not the same data and were never supposed to be; what must +// not change is the MAPPING between them. So the central test is a +// table: one row per (engine option shape × request shape), asserting +// both sides field by field. A rewrite that gets the recorded form right +// while pushing the wrong wire form fails it, and so does the reverse. +// +// The second half of the equivalence is the SSE race window. Its reader +// (`server/lib/mcode-acp.js`, ticket 08) is not this batch's to change, +// so this file imports it and pins the WRITER against it — including +// both reverse halves, because the failure this guards against is +// symmetric: a stamp written for a pick the user never made is as wrong +// as no stamp at all. +// +// Env is injected at MODULE scope, before any engine import, and the +// engine data dir points at a `mkTmpDir` fixture — the builtin tree +// `variantChannelFor` reads is the real file format, so a variant case +// that booted against the host's `~/.minimax/config.yaml` would be +// testing the operator's machine. + +import { test, describe, before, after, beforeEach } from "node:test"; +import assert from "node:assert/strict"; +import { readFileSync, writeFileSync } from "node:fs"; +import { fileURLToPath } from "node:url"; +import { join } from "node:path"; +import yaml from "js-yaml"; + +import { setupMocks, absPath, registerRpcMock } from "../../helpers/_setup.js"; +import { mkTmpDir, rmTmpDir } from "../../helpers/tmp.js"; + +// Pinned BEFORE any engine/engine-catalogue import below reads them. +const tmpBase = mkTmpDir("webui-model-writes-"); +process.env.MINIMAX_DATA_DIR = tmpBase; +process.env.MCODE_WEBUI_DATA_DIR = tmpBase; +process.env.MCODE_WEBUI_SETTINGS_PATH = `${tmpBase}/settings.json`; +process.env.MCODE_WEBUI_EVENTS_PATH = `${tmpBase}/events.jsonl`; +process.env.MCODE_WEBUI_SESSIONS_DB = `${tmpBase}/sessions.db`; +process.env.MCODE_WEBUI_UPLOAD_DIR = `${tmpBase}/uploads`; + +// The REAL permission mappers, captured at module scope — i.e. before +// any test context has registered the dispatch-through mock for +// `lib/mcode-rpc.js`. `setupMocks` replaces that module wholesale, so a +// mapping test that read `webuiPermissionToMcode` from it would be +// asserting the mock. This binding is the real function, and the mock's +// wrapper is pointed at it below. +const realRpc = await import(absPath("lib/mcode-rpc.js")); +const { variantChannelFor } = await import(absPath("lib/engine-catalogue.js")); + +/** Every name `engine/model-writes.js` exports. The namespace, not a subset. */ +const FACADE_EXPORTS = [ + "NO_SESSION_MODEL_WARNING", + "NO_SESSION_PERMISSION_WARNING", + "applyThinkingEffortMirror", + "modelSelectionTarget", + "planModelPickStamps", + "planModelSelectionPush", + "pushEngineModelSelection", + "pushEnginePermissionMode", + "resolveEngineModelConfigValue", + "resolvePermissionSelection", +]; + +let bust = 0; + +/** The exported names read out of the SOURCE, so a new one cannot slip past. */ +function exportedNamesOf(relative) { + const fileUrl = absPath(relative); + const src = readFileSync(fileURLToPath(fileUrl), "utf8"); + const names = new Set(); + for (const m of src.matchAll(/^export\s+(?:async\s+)?function\s+([A-Za-z_$][\w$]*)/gm)) names.add(m[1]); + for (const m of src.matchAll(/^export\s+(?:const|let|var|class)\s+([A-Za-z_$][\w$]*)/gm)) names.add(m[1]); + for (const m of src.matchAll(/^export\s*\{([^}]*)\}/gm)) { + for (const part of m[1].split(",")) { + const name = part.trim().split(/\s+as\s+/).pop().trim(); + if (name) names.add(name); + } + } + return names; +} + +// --------------------------------------------------------------------------- +// Fixtures — the engine's own shapes, transcribed +// --------------------------------------------------------------------------- + +/** A switchable builtin, advertised as the variant wire pair (control-state.ts#uniqueModelValues). */ +const VARIANT_MODEL_OPTION = { + type: "select", + id: "model", + name: "Model", + currentValue: "m:minimax_api:MiniMax-M3:v:thinking", + options: [ + { value: "m:minimax_api:MiniMax-M3:v:thinking", name: "MiniMax-M3 · thinking" }, + { value: "m:minimax_api:MiniMax-M3:v:none-thinking", name: "MiniMax-M3 · none-thinking" }, + { value: "m:minimax_api:MiniMax-M2.7:u", name: "MiniMax-M2.7" }, + ], +}; + +/** A forced_on builtin with an effort vocabulary: one bare wire form. */ +const EFFORT_MODEL_OPTION = { + type: "select", + id: "model", + name: "Model", + currentValue: "m:minimax_api:MiniMax-M3.1-Flash-Preview:u", + options: [ + { value: "m:minimax_api:MiniMax-M3.1-Flash-Preview:u", name: "M3.1-Flash-Preview" }, + ], +}; + +/** A custom-provider model: the provider segment is percent-encoded. */ +const CUSTOM_MODEL_OPTION = { + type: "select", + id: "model", + name: "Model", + currentValue: "m:custom_provider%3Azai-pro:glm-5.3:u", + options: [{ value: "m:custom_provider%3Azai-pro:glm-5.3:u", name: "glm-5.3" }], +}; + +/** The engine's materialised builtin tree, verbatim shapes. */ +const BUILTIN_MINIMAX_TREE = { + minimax: { + models: { + "MiniMax-M3": { + name: "MiniMax-M3", + reasoning: true, + thinking_config: { mode: "switchable", default_value: "true" }, + variants: { + "none-thinking": { thinking: { type: "disabled" } }, + thinking: { thinking: { type: "adaptive" } }, + }, + }, + "MiniMax-M3.1-Flash-Preview": { + name: "M3.1-Flash-Preview", + reasoning: true, + thinking_config: { mode: "forced_on" }, + thinking: { + effortOptions: ["default", "low", "medium", "high", "xhigh", "max"], + defaultEffort: "default", + }, + variants: { + "none-thinking": { thinking: { type: "disabled" } }, + thinking: { thinking: { type: "adaptive" } }, + }, + }, + "MiniMax-M2.7": { name: "MiniMax-M2.7", reasoning: true, thinking_config: { mode: "forced_on" } }, + }, + }, +}; + +/** + * Run a body with the engine's builtin tree on disk. The tree is written + * ONCE for the whole file (see `before`) and lives until `after` removes + * the tmp dir: `variantChannelFor` reads it on every call, so a + * write-then-delete wrapper would make the plan cases order-dependent for + * no gain — every case here wants the same tree. + */ +const withBuiltinTree = (body) => body(); + +/** A client state, shaped like the live one. */ +function fakeCs(overrides = {}) { + return { + model: { name: "minimax_api/MiniMax-M3", thinking: "", ...(overrides.model || {}) }, + permissions: "Full access", + ...(overrides.configOptions === undefined ? {} : { configOptions: overrides.configOptions }), + ...(overrides.mcodeSessionId === undefined ? {} : { mcodeSessionId: overrides.mcodeSessionId }), + }; +} + +/** Client state with a live session — the only shape that pushes. */ +function fakeCsWithSession(overrides = {}) { + return fakeCs({ mcodeSessionId: "mvs_b10_0000000000000000000000", ...overrides }); +} + +// --------------------------------------------------------------------------- +// Booting the facade +// --------------------------------------------------------------------------- + +/** + * Boot `engine/model-writes.js` with the standard webui mocks and an + * RPC recorder. `?bust=N` gives every boot its own module instance, which + * is what lets a test change a mock and still see the new one. + * + * The recorder is the ONLY thing the executor's engine half touches, so + * every assertion about "what the engine received" is a value read back + * from this list rather than a spy count. + */ +async function bootWithRpc(t, impl = {}) { + await setupMocks(t, {}); + const calls = []; + registerRpcMock({ + webuiPermissionToMcode: realRpc.webuiPermissionToMcode, + setConfigOption: async (sid, configId, value, cid) => { + calls.push({ sid, configId, value, cid }); + if (typeof impl.setConfigOption === "function") return impl.setConfigOption({ sid, configId, value, cid }); + return { ok: true, data: {} }; + }, + }); + const facade = await import(`${absPath("engine/model-writes.js")}?bust=${bust++}`); + return { facade, calls }; +} + +/** Boot without the RPC mock — for the pure derivations and the mappers. */ +async function bootPure(t) { + await setupMocks(t, {}); + return import(`${absPath("engine/model-writes.js")}?bust=${bust++}`); +} + +// The real SSE-race reader, imported with no mocks in play. Its module +// load may start the resident acp singleton, which `after` shuts down — +// the same teardown `test/lib/mcode-acp-ownership.check.mjs` documents. +let raceReader; +before(async () => { + raceReader = await import(absPath("lib/mcode-acp.js")); + // The engine's materialised builtin tree, in the real file format, for + // the whole run. Written AFTER the reader import so the reader is the + // real one with no fixture in its way. + writeFileSync(join(tmpBase, "config.yaml"), yaml.dump({ provider: BUILTIN_MINIMAX_TREE }), "utf8"); +}); +after(async () => { + const acp = await import(absPath("lib/acp-client.js")); + try { await acp.getMcodeAcpClient(); } catch { /* engine never started */ } + try { acp.shutdownMcodeAcpSingleton(); } catch { /* nothing to stop */ } + rmTmpDir(tmpBase); +}); + +beforeEach(() => { + registerRpcMock({ + setConfigOption: async () => ({ ok: true, data: {} }), + webuiPermissionToMcode: realRpc.webuiPermissionToMcode, + }); +}); + +// --------------------------------------------------------------------------- +// The export surface +// --------------------------------------------------------------------------- + +describe("model-writes facade — export surface", () => { + test("exports exactly the names this file pins, no more and no fewer", async () => { + const module = await import(absPath("engine/model-writes.js")); + const actual = Object.keys(module).filter((k) => k !== "default").sort(); + assert.deepEqual(actual, FACADE_EXPORTS); + }); + + test("the name list is derived from the SOURCE, so a new export cannot slip past the sweep", () => { + assert.deepEqual([...exportedNamesOf("engine/model-writes.js")].sort(), FACADE_EXPORTS); + }); + + test("routes/model.js imports the module DIRECTLY, and the facade index does not re-export it", () => { + // The same call B4 made for `model-reads.js`: this module statically + // reaches `lib/engine-catalogue.js` (and through it js-yaml), and + // `engine/index.js` is the one import site the whole server shares. + const route = readFileSync(fileURLToPath(absPath("routes/model.js")), "utf8"); + assert.match(route, /from "\.\.\/engine\/model-writes\.js"/); + const index = readFileSync(fileURLToPath(absPath("engine/index.js")), "utf8"); + assert.equal(index.includes("model-writes.js"), false, "engine/index.js must not pull it in"); + }); + + test("the RPC wrapper is reached through a dynamic import, never a static one", () => { + // The boot-path rule every engine family follows: mcode-rpc.js pulls + // the ACP client, and a static import here would put it on the + // import graph of anything that loads the facade. + const src = readFileSync(fileURLToPath(absPath("engine/model-writes.js")), "utf8"); + assert.equal(/^import .*mcode-rpc/m.test(src), false, "no static import of mcode-rpc.js"); + assert.match(src, /import\("\.\.\/lib\/mcode-rpc\.js"\)/, "the wrapper is reached through import()"); + }); +}); + +// --------------------------------------------------------------------------- +// The two warning strings — wire values, pinned +// --------------------------------------------------------------------------- + +describe("the no-session warnings", () => { + test("each endpoint keeps its OWN sentence, and they are not the same string", async (t) => { + const { NO_SESSION_MODEL_WARNING, NO_SESSION_PERMISSION_WARNING } = await bootPure(t); + assert.equal(NO_SESSION_MODEL_WARNING, "no mcode session yet — recorded for the next one"); + assert.equal(NO_SESSION_PERMISSION_WARNING, "no mcode session yet — applies to the next one"); + assert.notEqual(NO_SESSION_MODEL_WARNING, NO_SESSION_PERMISSION_WARNING); + }); +}); + +// --------------------------------------------------------------------------- +// modelSelectionTarget +// --------------------------------------------------------------------------- + +describe("modelSelectionTarget", () => { + test("the requested model wins; the recorded one is the fallback; empty is empty", async (t) => { + const { modelSelectionTarget } = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" } }); + assert.equal(modelSelectionTarget(cs, "zai-pro/glm-5.3"), "zai-pro/glm-5.3"); + assert.equal(modelSelectionTarget(cs, ""), "minimax_api/MiniMax-M3", "thinking-only update"); + assert.equal(modelSelectionTarget({ model: {} }, ""), ""); + assert.equal(modelSelectionTarget(null, ""), "", "a client state that is not there yet"); + }); +}); + +// --------------------------------------------------------------------------- +// resolveEngineModelConfigValue — the wire form +// --------------------------------------------------------------------------- + +describe("resolveEngineModelConfigValue — webui id → engine wire value", () => { + const cases = [ + { name: "bare builtin form", cs: { configOptions: [VARIANT_MODEL_OPTION] }, id: "minimax_api/MiniMax-M2.7", want: "m:minimax_api:MiniMax-M2.7:u" }, + { name: "custom provider, percent-encoded", cs: { configOptions: [CUSTOM_MODEL_OPTION] }, id: "zai-pro/glm-5.3", want: "m:custom_provider%3Azai-pro:glm-5.3:u" }, + { name: "a multi-segment model key", cs: { configOptions: [{ id: "model", options: [{ value: "m:nousresearch:deepseek/x:u", name: "deepseek/x" }] }] }, id: "nousresearch/deepseek/x", want: "m:nousresearch:deepseek/x:u" }, + { name: "a display name that differs from the model id", cs: { configOptions: [{ id: "model", options: [{ value: "m:p:deepseek/deepseek-v4.1-flash:u", name: "DeepSeek V4.1 Flash" }] }] }, id: "p/deepseek/deepseek-v4.1-flash", want: "m:p:deepseek/deepseek-v4.1-flash:u" }, + { name: "an id that is already the wire value", cs: { configOptions: [CUSTOM_MODEL_OPTION] }, id: "m:custom_provider%3Azai-pro:glm-5.3:u", want: "m:custom_provider%3Azai-pro:glm-5.3:u" }, + ]; + + for (const c of cases) { + test(`${c.name}`, async (t) => { + const { resolveEngineModelConfigValue } = await bootPure(t); + assert.equal(resolveEngineModelConfigValue(c.cs, c.id), c.want); + }); + } + + test("preferVariant narrows the deliberate ambiguity of a switchable builtin", async (t) => { + const { resolveEngineModelConfigValue } = await bootPure(t); + const cs = { configOptions: [VARIANT_MODEL_OPTION] }; + const id = "minimax_api/MiniMax-M3"; + // Both options carry the same bare name, so without a preference this + // is ambiguous and the answer is null — the caller then pushes the + // recorded id rather than picking the wrong variant. + assert.equal(resolveEngineModelConfigValue(cs, id), null); + assert.equal(resolveEngineModelConfigValue(cs, id, { preferVariant: "none-thinking" }), "m:minimax_api:MiniMax-M3:v:none-thinking"); + assert.equal(resolveEngineModelConfigValue(cs, id, { preferVariant: "thinking" }), "m:minimax_api:MiniMax-M3:v:thinking"); + // A variant the engine does not advertise falls through to the + // pre-ticket-36 outcome rather than inventing a wire form. + assert.equal(resolveEngineModelConfigValue(cs, id, { preferVariant: "nonsense" }), null); + }); + + test("null whenever the engine has not advertised the model yet, and for a junk id", async (t) => { + const { resolveEngineModelConfigValue } = await bootPure(t); + // No `model` option at all (the state before the first session event). + assert.equal(resolveEngineModelConfigValue({ configOptions: [] }, "minimax_api/MiniMax-M3"), null); + assert.equal(resolveEngineModelConfigValue({}, "minimax_api/MiniMax-M3"), null); + assert.equal(resolveEngineModelConfigValue({ configOptions: [VARIANT_MODEL_OPTION] }, ""), null); + assert.equal(resolveEngineModelConfigValue({ configOptions: [VARIANT_MODEL_OPTION] }, 42), null); + assert.equal(resolveEngineModelConfigValue({ configOptions: [VARIANT_MODEL_OPTION] }, null), null); + }); +}); + +// --------------------------------------------------------------------------- +// planModelSelectionPush — the two channels, field by field +// --------------------------------------------------------------------------- + +describe("planModelSelectionPush — the variant channel", () => { + test("model + thinking:'off' folds the level into ONE model push and drops the effort push", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" }, configOptions: [VARIANT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ + cs, + modelId: "minimax_api/MiniMax-M3", + thinkingWasProvided: true, + thinking: "off", + variantPlan: variantChannelFor("minimax_api/MiniMax-M3"), + }); + assert.equal(plan.channel, "variant"); + assert.equal(plan.target, "minimax_api/MiniMax-M3"); + assert.deepEqual(plan.modelPush, { value: "m:minimax_api:MiniMax-M3:v:none-thinking" }); + assert.equal(plan.thinkingPush, null); + assert.equal(plan.reportsModelSynced, true); + assert.equal(plan.carriedThinking, true); + }); + + test("a thinking-only update attaches the level to the RECORDED model", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3", thinking: "off" }, configOptions: [VARIANT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ + cs, + thinkingWasProvided: true, + thinking: "on", + variantPlan: variantChannelFor(cs.model.name), + }); + assert.equal(plan.modelPush.value, "m:minimax_api:MiniMax-M3:v:thinking"); + // No model in the request, so the model-sync field stays false even + // though the model push itself succeeded. Pre-B10 behaviour. + assert.equal(plan.reportsModelSynced, false); + assert.equal(plan.carriedThinking, true); + }); + + test("no user-chosen level falls back to the engine's DEFAULT variant and reports nothing synced", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3", thinking: "" }, configOptions: [VARIANT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ + cs, + modelId: "minimax_api/MiniMax-M3", + variantPlan: variantChannelFor("minimax_api/MiniMax-M3"), + }); + assert.equal(plan.modelPush.value, "m:minimax_api:MiniMax-M3:v:thinking", "default_value true → thinking"); + assert.equal(plan.carriedThinking, false, "an engine-default variant is not a sync"); + }); + + test("a CLEARED level is not a carried level, and means the ENGINE DEFAULT variant", async (t) => { + // The level normaliser only knows "on" and "off"; a cleared level is + // neither, so it lands on the engine's default variant — which is + // "thinking" for a `default_value: "true"` model. Reading the clear + // as "off" would be a plausible rewrite and a behaviour change, so + // it is pinned from both sides. + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3", thinking: "on" }, configOptions: [VARIANT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ + cs, + thinkingWasProvided: true, + thinking: "", + variantPlan: variantChannelFor(cs.model.name), + }); + assert.equal(plan.modelPush.value, "m:minimax_api:MiniMax-M3:v:thinking"); + assert.equal(plan.carriedThinking, false); + }); + + test("a MODEL-only pick carries a RECORDED level — the absent field is not the same as an empty one", async (t) => { + // The engine's own wire form is chosen from the variant the session + // already has, so a model-only pick on a switchable builtin still + // carries that level — and reports it as synced. Reading the absent + // field as "nothing carried" is a plausible rewrite that would flip + // `thinkingSynced` on a real, successful push. + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M2.7", thinking: "off" }, configOptions: [VARIANT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ + cs, + modelId: "minimax_api/MiniMax-M3", + variantPlan: variantChannelFor("minimax_api/MiniMax-M3"), + }); + assert.equal(plan.modelPush.value, "m:minimax_api:MiniMax-M3:v:none-thinking"); + assert.equal(plan.carriedThinking, true, "the recorded level rode the model push"); + }); + + test("a target the engine has not advertised falls back to the recorded id, not to null", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" }, configOptions: [] }); + const plan = facade.planModelSelectionPush({ + cs, + modelId: "minimax_api/MiniMax-M3", + thinkingWasProvided: true, + thinking: "off", + variantPlan: variantChannelFor("minimax_api/MiniMax-M3"), + }); + assert.equal(plan.modelPush.value, "minimax_api/MiniMax-M3", "the raw webui id, unchanged"); + }); + + test("a NON-builtin id never rides the variant channel", async (t) => { + const facade = await bootPure(t); + // `variantChannelFor` returns null for anything outside the builtin + // provider, even a model that shares a bare name with a builtin. + assert.equal(variantChannelFor("zai-pro/glm-5.3"), null); + const cs = fakeCs({ model: { name: "zai-pro/glm-5.3" }, configOptions: [CUSTOM_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ cs, modelId: "zai-pro/glm-5.3", thinkingWasProvided: true, thinking: "high" }); + assert.equal(plan.channel, "effort"); + }); +}); + +describe("planModelSelectionPush — the effort channel", () => { + test("model first, then effort — the engine's own contract", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M2.7" }, configOptions: [EFFORT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ cs, modelId: "minimax_api/MiniMax-M3.1-Flash-Preview", thinkingWasProvided: true, thinking: "high" }); + assert.equal(plan.channel, "effort"); + assert.deepEqual(plan.modelPush, { value: "m:minimax_api:MiniMax-M3.1-Flash-Preview:u" }); + assert.deepEqual(plan.thinkingPush, { value: "high" }); + assert.equal(plan.reportsModelSynced, true); + assert.equal(plan.carriedThinking, true); + }); + + test("a thinking-only update pushes ONLY the effort, and reports no model sync", async (t) => { + // The asymmetry with the variant channel, and it is pre-existing: a + // non-switchable model has no wire form that carries an effort, so + // the level travels on its own config id. `mcodeSynced` stays false + // because no model was in the request. + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3.1-Flash-Preview" }, configOptions: [EFFORT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ cs, thinkingWasProvided: true, thinking: "high" }); + assert.equal(plan.modelPush, null, "no model push — nothing in the request names a model"); + assert.deepEqual(plan.thinkingPush, { value: "high" }); + assert.equal(plan.reportsModelSynced, false); + assert.equal(plan.carriedThinking, true); + }); + + test("a cleared level plans no effort push, and the model id is left alone", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3.1-Flash-Preview", thinking: "high" }, configOptions: [EFFORT_MODEL_OPTION] }); + const plan = facade.planModelSelectionPush({ cs, thinkingWasProvided: true, thinking: "" }); + assert.equal(plan.modelPush, null); + assert.equal(plan.thinkingPush, null); + assert.equal(plan.carriedThinking, false); + }); + + test("an absent thinking field is NOT the same as a cleared one", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3.1-Flash-Preview", thinking: "high" }, configOptions: [EFFORT_MODEL_OPTION] }); + const absent = facade.planModelSelectionPush({ cs, modelId: "minimax_api/MiniMax-M3.1-Flash-Preview" }); + assert.equal(absent.thinkingPush, null, "the recorded level is not pushed on a model-only pick"); + assert.equal(absent.carriedThinking, false); + }); + + test("the model value falls back to the recorded id when the engine advertises nothing yet", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M2.7" }, configOptions: [] }); + const plan = facade.planModelSelectionPush({ cs, modelId: "minimax_api/MiniMax-M2.7" }); + assert.deepEqual(plan.modelPush, { value: "minimax_api/MiniMax-M2.7" }); + }); +}); + +// --------------------------------------------------------------------------- +// THE TWO FORMS, FIELD BY FIELD +// --------------------------------------------------------------------------- + +describe("wire form ↔ recorded selection — field by field", () => { + /** + * One row per (engine option shape × request shape). `wire` is what the + * engine must receive, in order; `recorded` is what webui must keep; the + * rest is the response triple. Every field is asserted on every row — + * a row that only checked the wire form would let the recorded form rot + * while the suite stayed green, which is the half of the equivalence + * nobody looks at. + */ + const ROWS = [ + { + name: "bare builtin, model only", + channel: "effort", + configOptions: [VARIANT_MODEL_OPTION], + cs: { model: { name: "minimax_api/MiniMax-M2.7", thinking: "" } }, + request: { model: "minimax_api/MiniMax-M2.7" }, + wire: [{ configId: "model", value: "m:minimax_api:MiniMax-M2.7:u" }], + recorded: { name: "minimax_api/MiniMax-M2.7", thinking: "" }, + result: { mcodeSynced: true, thinkingSynced: false, warning: null }, + }, + { + name: "bare builtin, model + effort", + channel: "effort", + configOptions: [EFFORT_MODEL_OPTION], + cs: { model: { name: "minimax_api/MiniMax-M3.1-Flash-Preview", thinking: "" } }, + request: { model: "minimax_api/MiniMax-M3.1-Flash-Preview", thinking: "high" }, + wire: [ + { configId: "model", value: "m:minimax_api:MiniMax-M3.1-Flash-Preview:u" }, + { configId: "thinkingEffort", value: "high" }, + ], + recorded: { name: "minimax_api/MiniMax-M3.1-Flash-Preview", thinking: "high" }, + result: { mcodeSynced: true, thinkingSynced: true, warning: null }, + }, + { + name: "custom provider, model + effort", + channel: "effort", + configOptions: [CUSTOM_MODEL_OPTION], + cs: { model: { name: "zai-pro/glm-5.3", thinking: "" } }, + request: { model: "zai-pro/glm-5.3", thinking: "high" }, + wire: [ + { configId: "model", value: "m:custom_provider%3Azai-pro:glm-5.3:u" }, + { configId: "thinkingEffort", value: "high" }, + ], + recorded: { name: "zai-pro/glm-5.3", thinking: "high" }, + result: { mcodeSynced: true, thinkingSynced: true, warning: null }, + }, + { + name: "switchable builtin, model + level off — ONE push", + channel: "variant", + configOptions: [VARIANT_MODEL_OPTION], + cs: { model: { name: "minimax_api/MiniMax-M3", thinking: "" } }, + request: { model: "minimax_api/MiniMax-M3", thinking: "off" }, + wire: [{ configId: "model", value: "m:minimax_api:MiniMax-M3:v:none-thinking" }], + recorded: { name: "minimax_api/MiniMax-M3", thinking: "off" }, + result: { mcodeSynced: true, thinkingSynced: true, warning: null }, + }, + { + name: "switchable builtin, level only — the recorded model carries it", + channel: "variant", + configOptions: [VARIANT_MODEL_OPTION], + cs: { model: { name: "minimax_api/MiniMax-M3", thinking: "off" } }, + request: { thinking: "on" }, + wire: [{ configId: "model", value: "m:minimax_api:MiniMax-M3:v:thinking" }], + recorded: { name: "minimax_api/MiniMax-M3", thinking: "on" }, + // No model in the request, so the model-sync field is false even + // though the model push succeeded. Pinned because it reads like a + // bug and is not one. + result: { mcodeSynced: false, thinkingSynced: true, warning: null }, + }, + { + name: "switchable builtin, model only — engine default variant, nothing synced", + channel: "variant", + configOptions: [VARIANT_MODEL_OPTION], + cs: { model: { name: "minimax_api/MiniMax-M3", thinking: "" } }, + request: { model: "minimax_api/MiniMax-M3" }, + wire: [{ configId: "model", value: "m:minimax_api:MiniMax-M3:v:thinking" }], + recorded: { name: "minimax_api/MiniMax-M3", thinking: "" }, + result: { mcodeSynced: true, thinkingSynced: false, warning: null }, + }, + { + name: "switchable builtin, model only, recorded level carried", + channel: "variant", + configOptions: [VARIANT_MODEL_OPTION], + cs: { model: { name: "minimax_api/MiniMax-M2.7", thinking: "off" } }, + request: { model: "minimax_api/MiniMax-M3" }, + wire: [{ configId: "model", value: "m:minimax_api:MiniMax-M3:v:none-thinking" }], + recorded: { name: "minimax_api/MiniMax-M3", thinking: "off" }, + result: { mcodeSynced: true, thinkingSynced: true, warning: null }, + }, + { + name: "engine advertises no model option yet — the recorded id goes out verbatim", + channel: "variant", + configOptions: [], + cs: { model: { name: "minimax_api/MiniMax-M2.7", thinking: "" } }, + request: { model: "minimax_api/MiniMax-M3" }, + wire: [{ configId: "model", value: "minimax_api/MiniMax-M3" }], + recorded: { name: "minimax_api/MiniMax-M3", thinking: "" }, + result: { mcodeSynced: true, thinkingSynced: false, warning: null }, + }, + ]; + + for (const row of ROWS) { + test(row.name, async (t) => { + const { facade, calls } = await bootWithRpc(t); + const cs = fakeCsWithSession({ model: row.cs.model, configOptions: row.configOptions }); + await withBuiltinTree(async () => { + // The row's `request` is the route's BODY; this is the route's own + // translation of it, so the recorded-form half of the row is a + // simulation of `handleSetModel` rather than a second truth. + const body = { ...row.request }; + const modelId = typeof body.model === "string" ? body.model.trim() : ""; + const thinkingWasProvided = Object.prototype.hasOwnProperty.call(body, "thinking"); + const thinking = thinkingWasProvided ? body.thinking.trim() : undefined; + if (modelId) cs.model.name = modelId; + if (thinkingWasProvided) cs.model.thinking = thinking; + const r = await facade.pushEngineModelSelection({ cs, cid: "cid-b10", modelId, thinkingWasProvided, thinking }); + // --- the wire form, in order ------------------------------------- + assert.deepEqual( + calls.map((c) => ({ configId: c.configId, value: c.value })), + row.wire, + "what the engine receives", + ); + assert.ok(calls.every((c) => c.sid === "mvs_b10_0000000000000000000000")); + assert.ok(calls.every((c) => c.cid === "cid-b10"), "the cid reaches the wrapper on every push"); + // --- the response triple ------------------------------------------ + assert.equal(r.mcodeSynced, row.result.mcodeSynced, "mcodeSynced"); + assert.equal(r.thinkingSynced, row.result.thinkingSynced, "thinkingSynced"); + assert.equal(r.warning, row.result.warning, "warning"); + assert.equal(r.channel, row.channel, "which channel the plan took"); + // --- the recorded form (the route's writes, above) ---------------- + assert.equal(cs.model.name, row.recorded.name, "cs.model.name"); + assert.equal(cs.model.thinking, row.recorded.thinking, "cs.model.thinking"); + }); + }); + } + + test("a rejected model push surfaces the engine's error and never invents a sync", async (t) => { + const { facade, calls } = await bootWithRpc(t, { + setConfigOption: ({ configId }) => + configId === "model" + ? { ok: false, error: "unknown model" } + : { ok: true, data: {} }, + }); + const cs = fakeCsWithSession({ model: { name: "minimax_api/MiniMax-M2.7" }, configOptions: [EFFORT_MODEL_OPTION] }); + await withBuiltinTree(async () => { + const r = await facade.pushEngineModelSelection({ + cs, + cid: "cid-b10", + modelId: "minimax_api/MiniMax-M3.1-Flash-Preview", + thinkingWasProvided: true, + thinking: "high", + }); + assert.equal(r.mcodeSynced, false); + assert.equal(r.thinkingSynced, true, "the effort push was accepted on its own"); + // The MODEL rejection stays the warning; the accepted effort does + // not overwrite it. + assert.equal(r.warning, "unknown model"); + assert.equal(calls.length, 2); + }); + }); + + test("a REJECTED variant push reports the engine's error and claims nothing synced", async (t) => { + const { facade, calls } = await bootWithRpc(t, { setConfigOption: () => ({ ok: false, error: "variant refused" }) }); + const cs = fakeCsWithSession({ model: { name: "minimax_api/MiniMax-M3" }, configOptions: [VARIANT_MODEL_OPTION] }); + const r = await facade.pushEngineModelSelection({ cs, cid: "cid-b10", modelId: "minimax_api/MiniMax-M3", thinkingWasProvided: true, thinking: "off" }); + assert.equal(r.channel, "variant"); + assert.equal(r.mcodeSynced, false); + assert.equal(r.thinkingSynced, false, "a rejected push synced nothing, not even the level"); + assert.equal(r.warning, "variant refused"); + assert.equal(r.thinkingMirror, null, "the variant channel never mirrors the effort option"); + assert.equal(calls.length, 1, "and it is still ONE push"); + }); + + test("a rejected effort push takes the warning only when the model push did not take it", async (t) => { + const { facade } = await bootWithRpc(t, { + setConfigOption: ({ configId }) => + configId === "thinkingEffort" ? { ok: false, error: "effort refused" } : { ok: true, data: {} }, + }); + const cs = fakeCsWithSession({ model: { name: "minimax_api/MiniMax-M2.7" }, configOptions: [EFFORT_MODEL_OPTION] }); + await withBuiltinTree(async () => { + const r = await facade.pushEngineModelSelection({ + cs, + cid: "cid-b10", + modelId: "minimax_api/MiniMax-M3.1-Flash-Preview", + thinkingWasProvided: true, + thinking: "high", + }); + assert.equal(r.mcodeSynced, true); + assert.equal(r.thinkingSynced, false); + assert.equal(r.warning, "effort refused"); + }); + }); + + test("no session: nothing is pushed, and the warning is this endpoint's own sentence", async (t) => { + const { facade, calls } = await bootWithRpc(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M2.7" }, configOptions: [VARIANT_MODEL_OPTION] }); + const r = await facade.pushEngineModelSelection({ cs, cid: "cid-b10", modelId: "minimax_api/MiniMax-M2.7" }); + assert.equal(r.channel, "no-session"); + assert.equal(r.mcodeSynced, false); + assert.equal(r.thinkingSynced, false); + assert.equal(r.thinkingMirror, null); + assert.equal(r.warning, facade.NO_SESSION_MODEL_WARNING); + assert.deepEqual(calls, []); + }); +}); + +// --------------------------------------------------------------------------- +// The SSE 4s race window — the writer (this batch) against the real reader +// --------------------------------------------------------------------------- + +describe("the SSE race window — writer and reader, pinned together", () => { + const FRESH = () => Date.now(); + + test("a pick inside the window defers the engine mirror, for every field it carried", async (t) => { + const facade = await bootPure(t); + const pickAt = FRESH(); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" } }); + Object.assign(cs.model, facade.planModelPickStamps({ modelId: "minimax_api/MiniMax-M3", thinkingWasProvided: true, contextWindowWasProvided: true }, pickAt)); + const engine = { currentValue: "m:minimax_api:MiniMax-M3:v:thinking" }; + assert.equal(raceReader.shouldMirrorToModelName(cs, engine), false, "model mirror deferred"); + assert.equal(raceReader.shouldMirrorToThinkingField(cs), false, "thinking mirror deferred"); + }); + + test("REVERSE HALF — a pick OLDER than the window does not defer: engine truth wins again", async (t) => { + const facade = await bootPure(t); + // A cross-client pick that landed 10s ago, or a sluggish engine + // answering late, must be allowed to catch up. This is the half a + // "always stamp" simplification breaks. + const pickAt = Date.now() - raceReader.PICK_DEFER_WINDOW_MS - 1000; + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" } }); + Object.assign(cs.model, facade.planModelPickStamps({ modelId: "minimax_api/MiniMax-M3", thinkingWasProvided: true }, pickAt)); + const engine = { currentValue: "m:minimax_api:MiniMax-M3:v:thinking" }; + assert.equal(raceReader.shouldMirrorToModelName(cs, engine), true, "model mirror reactivated"); + assert.equal(raceReader.shouldMirrorToThinkingField(cs), true, "thinking mirror reactivated"); + }); + + test("REVERSE HALF — a field the request did NOT carry keeps its old stamp and mirrors immediately", async (t) => { + const facade = await bootPure(t); + // A thinking-only pick must not suppress a later cross-client MODEL + // change. Stamping every field would, and the per-field independence + // that ticket 08 bought would be lost with it. + const pickAt = FRESH(); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" } }); + Object.assign(cs.model, facade.planModelPickStamps({ thinkingWasProvided: true }, pickAt)); + assert.equal(cs.model.modelPickedAt, undefined, "the model field was not stamped"); + assert.equal(raceReader.shouldMirrorToModelName(cs, { currentValue: "m:z:p:u" }), true, "model mirror NOT deferred"); + assert.equal(raceReader.shouldMirrorToThinkingField(cs), false, "thinking mirror deferred"); + }); + + test("REVERSE HALF — a context-window-only pick defers nothing the reader looks at", async (t) => { + const facade = await bootPure(t); + const cs = fakeCs({ model: { name: "minimax_api/MiniMax-M3" } }); + Object.assign(cs.model, facade.planModelPickStamps({ contextWindowWasProvided: true }, FRESH())); + assert.equal(raceReader.shouldMirrorToModelName(cs, { currentValue: "m:z:p:u" }), true); + assert.equal(raceReader.shouldMirrorToThinkingField(cs), true); + }); + + test("ONE timestamp for the whole request — the window is a race window, not three", async (t) => { + const facade = await bootPure(t); + const pickAt = 1_700_000_000_000; + const stamps = facade.planModelPickStamps({ modelId: "minimax_api/MiniMax-M3", thinkingWasProvided: true, contextWindowWasProvided: true }, pickAt); + assert.deepEqual(stamps, { + modelPickedAt: pickAt, + thinkingPickedAt: pickAt, + contextWindowPickedAt: pickAt, + }); + assert.equal(new Set(Object.values(stamps)).size, 1, "all three fields share one instant"); + }); + + test("a request that carried nothing stamps nothing", async (t) => { + const facade = await bootPure(t); + assert.deepEqual(facade.planModelPickStamps({}, Date.now()), {}); + assert.deepEqual(facade.planModelPickStamps({ modelId: "" }, Date.now()), {}, "an empty model id is not a pick"); + }); +}); + +// --------------------------------------------------------------------------- +// applyThinkingEffortMirror — the local snapshot rule +// --------------------------------------------------------------------------- + +describe("applyThinkingEffortMirror", () => { + test("set claims the engine's new value; clear DROPS it; null does nothing", async (t) => { + const { applyThinkingEffortMirror } = await bootPure(t); + const opts = [ + { id: "model", currentValue: "m:minimax_api:MiniMax-M3:u" }, + { id: "thinkingEffort", currentValue: "low" }, + ]; + assert.equal(applyThinkingEffortMirror(opts, { kind: "set", value: "high" }), 1); + assert.equal(opts[1].currentValue, "high"); + assert.equal(opts[0].currentValue, "m:minimax_api:MiniMax-M3:u", "the model option is untouched"); + assert.equal(applyThinkingEffortMirror(opts, { kind: "clear" }), 1); + assert.equal("currentValue" in opts[1], false, "cleared, not emptied — an empty string means the engine's default"); + assert.equal(applyThinkingEffortMirror(opts, null), 0); + }); + + test("a snapshot with no thinkingEffort option is a no-op, and says so", async (t) => { + const { applyThinkingEffortMirror } = await bootPure(t); + assert.equal(applyThinkingEffortMirror([{ id: "model" }], { kind: "set", value: "high" }), 0); + assert.equal(applyThinkingEffortMirror(undefined, { kind: "set", value: "high" }), 0); + assert.equal(applyThinkingEffortMirror(null, { kind: "clear" }), 0); + assert.equal(applyThinkingEffortMirror("not an array", { kind: "clear" }), 0); + }); + + test("the route applies exactly the mirror the executor asked for", async (t) => { + const { facade, calls } = await bootWithRpc(t); + const cs = fakeCsWithSession({ + model: { name: "minimax_api/MiniMax-M2.7" }, + configOptions: [{ id: "thinkingEffort", currentValue: "low" }, { id: "model" }], + }); + await withBuiltinTree(async () => { + const r = await facade.pushEngineModelSelection({ cs, cid: "c", modelId: "minimax_api/MiniMax-M2.7", thinkingWasProvided: true, thinking: "high" }); + assert.deepEqual(r.thinkingMirror, { kind: "set", value: "high" }); + facade.applyThinkingEffortMirror(cs.configOptions, r.thinkingMirror); + assert.equal(cs.configOptions[0].currentValue, "high"); + // A model-only pick changes nothing here: the engine has not + // reported a new effort, so the mirror must not invent one. + const other = fakeCsWithSession({ + model: { name: "minimax_api/MiniMax-M2.7" }, + configOptions: [{ id: "thinkingEffort", currentValue: "low" }], + }); + const r2 = await facade.pushEngineModelSelection({ cs: other, cid: "c", modelId: "minimax_api/MiniMax-M2.7" }); + assert.equal(r2.thinkingMirror, null); + assert.equal(other.configOptions[0].currentValue, "low", "untouched"); + }); + assert.ok(calls.length >= 1); + }); + + test("a REFUSED effort push mirrors nothing, and a cleared one drops the value", async (t) => { + const { facade } = await bootWithRpc(t, { + setConfigOption: ({ configId }) => + configId === "thinkingEffort" ? { ok: false, error: "no" } : { ok: true, data: {} }, + }); + const cs = fakeCsWithSession({ + model: { name: "minimax_api/MiniMax-M2.7" }, + configOptions: [{ id: "thinkingEffort", currentValue: "low" }], + }); + await withBuiltinTree(async () => { + const refused = await facade.pushEngineModelSelection({ cs, cid: "c", modelId: "minimax_api/MiniMax-M2.7", thinkingWasProvided: true, thinking: "high" }); + assert.equal(refused.thinkingMirror, null, "an unaccepted push is not mirrored"); + assert.equal(cs.configOptions[0].currentValue, "low", "untouched"); + // Model changed AND effort cleared → the mirror is dropped, and it + // does NOT depend on the model push having succeeded. + const cleared = await facade.pushEngineModelSelection({ cs, cid: "c", modelId: "minimax_api/MiniMax-M2.7", thinkingWasProvided: true, thinking: "" }); + assert.deepEqual(cleared.thinkingMirror, { kind: "clear" }); + facade.applyThinkingEffortMirror(cs.configOptions, cleared.thinkingMirror); + assert.equal("currentValue" in cs.configOptions[0], false); + }); + }); + + test("a clear with NO model change mirrors nothing — the next config_option_update reports it", async (t) => { + const { facade } = await bootWithRpc(t); + const cs = fakeCsWithSession({ + model: { name: "minimax_api/MiniMax-M2.7" }, + configOptions: [{ id: "thinkingEffort", currentValue: "low" }], + }); + await withBuiltinTree(async () => { + const r = await facade.pushEngineModelSelection({ cs, cid: "c", thinkingWasProvided: true, thinking: "" }); + assert.equal(r.thinkingMirror, null); + assert.equal(cs.configOptions[0].currentValue, "low", "untouched"); + }); + }); +}); + +// --------------------------------------------------------------------------- +// resolvePermissionSelection — both forms of one mode +// --------------------------------------------------------------------------- + +describe("resolvePermissionSelection — the two forms of one mode", () => { + const TABLE = [ + { mode: "ask", label: "Ask", mcodeValue: "default" }, + { mode: "auto", label: "Auto", mcodeValue: "auto" }, + { mode: "read", label: "Read", mcodeValue: "read" }, + { mode: "off", label: "Off", mcodeValue: "off" }, + { mode: "full", label: "Full access", mcodeValue: "bypassPermissions" }, + ]; + + for (const row of TABLE) { + test(`${row.mode} → label "${row.label}", engine value "${row.mcodeValue}"`, async (t) => { + const { resolvePermissionSelection } = await bootPure(t); + assert.deepEqual(await resolvePermissionSelection(row.mode), { label: row.label, mcodeValue: row.mcodeValue }); + }); + } + + test("case is folded, and an unknown mode still gets a label", async (t) => { + const { resolvePermissionSelection } = await bootPure(t); + assert.deepEqual(await resolvePermissionSelection("FULL"), { label: "Full access", mcodeValue: "bypassPermissions" }); + assert.deepEqual(await resolvePermissionSelection("AsK"), { label: "Ask", mcodeValue: "default" }); + }); + + test("an unknown or missing mode has NO engine value — the two mappers disagree on purpose", async (t) => { + // The label mapper falls back to `full` so the UI always has + // something to show; the engine mapper returns null because there is + // no engine word for a mode the user invented. So the endpoint + // records a label and does NOT push — which is exactly why + // `pushEnginePermissionMode` guards on the value and not only on the + // session. Pinned as a value because "unknown → Full access" reads + // like it should also push `bypassPermissions`, and it must not. + const { resolvePermissionSelection } = await bootPure(t); + // Only a NON-EMPTY unknown string: `""` and `undefined` are falsy and + // fold to the `full` default before either mapper sees them. + for (const mode of ["nonsense", "nope", "Ask!"]) { + assert.deepEqual( + await resolvePermissionSelection(mode), + { label: "Full access", mcodeValue: null }, + JSON.stringify(mode), + ); + } + for (const mode of ["", undefined, null]) { + assert.deepEqual( + await resolvePermissionSelection(mode), + { label: "Full access", mcodeValue: "bypassPermissions" }, + JSON.stringify(mode), + ); + } + }); + + test("label and engine value come from TWO different mappers, and both are in play", async (t) => { + // The seam's whole reason: a fifth form added to one mapper and not + // the other would be a mode webui records and never delivers. + const { resolvePermissionSelection } = await bootPure(t); + const { webuiModeToLabel } = await import(absPath("lib/interaction/permission-presets.js")); + for (const row of TABLE) { + const got = await resolvePermissionSelection(row.mode); + assert.equal(got.label, webuiModeToLabel(row.mode), "label from permission-presets"); + assert.equal(got.mcodeValue, realRpc.webuiPermissionToMcode(row.mode), "engine value from mcode-rpc"); + } + }); +}); + +// --------------------------------------------------------------------------- +// pushEnginePermissionMode +// --------------------------------------------------------------------------- + +describe("pushEnginePermissionMode", () => { + test("a live session gets exactly one permissionMode push, on this cid", async (t) => { + const { facade, calls } = await bootWithRpc(t); + const cs = fakeCsWithSession(); + const r = await facade.pushEnginePermissionMode({ cs, cid: "cid-b10", mcodeValue: "default" }); + assert.deepEqual(calls, [{ sid: "mvs_b10_0000000000000000000000", configId: "permissionMode", value: "default", cid: "cid-b10" }]); + assert.deepEqual(r, { mcodeSynced: true, warning: null }); + }); + + test("no session: local only, and THIS endpoint's warning sentence", async (t) => { + const { facade, calls } = await bootWithRpc(t); + const cs = fakeCs(); + const r = await facade.pushEnginePermissionMode({ cs, cid: "cid-b10", mcodeValue: "default" }); + assert.deepEqual(calls, []); + assert.equal(r.mcodeSynced, false); + assert.equal(r.warning, facade.NO_SESSION_PERMISSION_WARNING); + }); + + test("a mode with no engine value is recorded but not pushed, and claims no sync", async (t) => { + // Not a shape today's mapper produces; the guard is the difference + // between "the engine is in this mode" and "we hope it is". + const { facade, calls } = await bootWithRpc(t); + const cs = fakeCsWithSession(); + const r = await facade.pushEnginePermissionMode({ cs, cid: "cid-b10", mcodeValue: null }); + assert.deepEqual(calls, []); + assert.equal(r.mcodeSynced, false); + assert.equal(r.warning, null, "nothing went wrong; there was simply nothing to say"); + }); + + test("a rejected push surfaces the engine's error verbatim", async (t) => { + const { facade } = await bootWithRpc(t, { setConfigOption: () => ({ ok: false, error: "mode refused" }) }); + const cs = fakeCsWithSession(); + const r = await facade.pushEnginePermissionMode({ cs, cid: "cid-b10", mcodeValue: "auto" }); + assert.equal(r.mcodeSynced, false); + assert.equal(r.warning, "mode refused"); + }); +}); diff --git a/release/public-source.json b/release/public-source.json index 7421b7eb4..456b56628 100644 --- a/release/public-source.json +++ b/release/public-source.json @@ -2698,6 +2698,7 @@ "packages/local-runtime/src/website-management/deployed-website-source.ts", "packages/local-runtime/src/website-management/local-website-management-client.ts", "packages/local-runtime/src/worktrees/branch.ts", + "packages/local-runtime/test/unit/agent-name-conflict-migration-lock.test.ts", "packages/local-runtime/test/unit/background-task-test-helpers.ts", "packages/local-runtime/test/unit/child-bash-lifecycle.test.ts", "packages/local-runtime/test/unit/content-safety-api-fail-policy.test.ts", @@ -3455,6 +3456,7 @@ "packages/webui/server/engine/interrupt.js", "packages/webui/server/engine/mode-writes.js", "packages/webui/server/engine/model-reads.js", + "packages/webui/server/engine/model-writes.js", "packages/webui/server/engine/providers/local-runtime-v2.capabilities.js", "packages/webui/server/engine/providers/local-runtime-v2.js", "packages/webui/server/engine/providers/tui-runtime-adapter.js", @@ -3610,6 +3612,7 @@ "packages/webui/test/lib/engine/interrupt.test.js", "packages/webui/test/lib/engine/mode-writes.test.js", "packages/webui/test/lib/engine/model-reads.test.js", + "packages/webui/test/lib/engine/model-writes.test.js", "packages/webui/test/lib/engine/session-export.test.js", "packages/webui/test/lib/engine/session-load.test.js", "packages/webui/test/lib/engine/session-reads.test.js", diff --git a/scripts/test-tmp-leak.check.mjs b/scripts/test-tmp-leak.check.mjs index 6fa9d67b9..0520ba7da 100644 --- a/scripts/test-tmp-leak.check.mjs +++ b/scripts/test-tmp-leak.check.mjs @@ -290,6 +290,7 @@ const KNOWN_PREFIXES = [ "webui-model-engine-cat-", "webui-model-reads-", "webui-model-user-level-", + "webui-model-writes-", "webui-models-merge-", "webui-origingate-events-", "webui-origingate-settings-", diff --git a/test/vitest-suites.json b/test/vitest-suites.json index dce989777..6ef1ba3a6 100644 --- a/test/vitest-suites.json +++ b/test/vitest-suites.json @@ -81,6 +81,7 @@ "packages/local-runtime-v2/test/unit/agent/storage/agent.repository.test.ts", "packages/local-runtime-v2/test/unit/agent/storage/canonical-agent-config.test.ts", "packages/local-runtime-v2/test/unit/compat/v1/runtime.test.ts", + "packages/local-runtime/test/unit/agent-name-conflict-migration-lock.test.ts", "packages/local-runtime/test/unit/child-bash-lifecycle.test.ts", "packages/local-runtime/test/unit/content-safety-api-fail-policy.test.ts", "packages/local-runtime/test/unit/content-safety-api-v2.test.ts",