Skip to content

Latest commit

 

History

History
399 lines (319 loc) · 29.1 KB

File metadata and controls

399 lines (319 loc) · 29.1 KB

TUI capability coverage

The current capability target is TUI 0.4.12; see version and evidence baseline for the separate workspace and embedded-tool versions. “Restored” below describes implementation and assembly, not acceptance of every account or online service.

Regular-mode terminal history

Regular mode preserves native terminal history, including shell output from before MCode started. Rows already emitted into scrollback are snapshots of what was shown at that time; folding, completed tools and later edits do not rewrite those rows. The live application view remains the source for the current transcript.

Long turns keep a bounded live projection. Dropping its oldest cells does not replay the welcome banner or erase earlier output. When activity or Todo panels shrink, released rows temporarily remain blank between the transcript and footer; new output consumes that space. Tables and code blocks remain contiguous.

When an unrelated transcript replaces the current projection, a labelled refresh boundary separates the retained output from the new document. Resize and forced redraw preserve pending output and native history. Explicit resume-start history clearing remains a separate startup policy. Native image history is retained until the terminal itself evicts it.

Transient overlays use an alternate buffer so opening, nesting or resizing panels cannot place menu rows in native history. Closing the last overlay restores the main buffer and reconciles background output; terminal shutdown also restores it.

Output speed

The activity line and completed-turn summary show provider output tokens divided by model generation time. Timing starts at the first nonempty text, reasoning, or tool token and ends when the model finishes, excluding first-token wait and tool execution. Multiple responses in a turn use total tokens divided by total generation time.

Speed appears after the first response with both provider usage and generation timing; streaming text is not estimated. Confirmed speed stays visible during later requests and tools, duplicate messages do not count twice, and a new turn resets the samples. Messages without generation timing are excluded. Values below 10 tok/s use one decimal place; higher values are rounded to integers. Batched provider events measure the observed generation window, not server hardware throughput.

Terminal titles and notifications

Terminal titles show the current state, session name and MCode, for example Needs approval | Fix login | MCode. Renaming or switching a session updates the title. Unnamed sessions use the project name and a short session ID. Titles are cleared when MCode exits or suspends and reapplied when it resumes.

Configure these presentation settings in the MCode data directory's config.yaml:

tui:
  terminalTitle: [status, session-name, app-name]
  notifications:
    when: unfocused
    method: auto
    events: [turn-complete, turn-failed, permission-required, question-required]

Title items can be ordered or omitted; project-name is also available. Set terminalTitle to null or [] to disable title updates. Unknown items are ignored. Notification when accepts unfocused, always or never; method accepts auto, osc9, osc777 or bel. Omitting events enables all four events; [] disables them. Apply configuration changes by restarting MCode.

Notifications identify the session and suppress duplicates. Completion waits for the session's queue to finish; failed turns and requests for input can notify independently. Known foreground focus suppresses notifications by default. When focus is unknown, delivery is best-effort; cmux manages its own surface focus. Automatic delivery uses the detected terminal's notification protocol or falls back to a bell. The existing Windows toast bridge is restricted to local Windows or WSL interop. Terminal settings and OS notification permissions still apply.

VS Code normally displays a process name in its terminal tabs. To display MCode's session titles, use this VS Code setting:

"terminal.integrated.tabs.title": "${sequence}"

A manually assigned tab title overrides automatic titles. VS Code's bell is a terminal-tab indicator, not a guarantee of a desktop notification. See the VS Code terminal appearance documentation. Inside tmux, OSC notifications require passthrough and support from the outer terminal; use method: bel for a bell fallback.

The evidence column summarizes the historical TUI 0.3.11 restoration record from 2026-09-11. It does not claim fresh TUI 0.4.12 live-service acceptance. Use current verification status for checks run against the updated source and explicit NOT RUN boundaries.

Capability matrix

Status legend: ✅ supported · ⚠ partial · ❌ unsupported · 🚧 requires engine-side support.

Capability Implementation Evidence
MiniMax login, logout, Token Plan, quota, check-in OAuth Core, account clients, and TUI / CLI entry points restored Login-state, auth-command, check-in, provider, and ACP tests; production Token Plan session and resume passed
BYOK / custom models Retained; headless model overrides no longer inherit the default Token Plan login requirement Local protocol server drives runtime, resume, and file tools; live BYOK session passed; managed models still reject unauthenticated requests
mcode-tools Public package's unchanged 0.0.4 artifact, launcher, and short-lived token leases restored Archive SHA-512, CLI SHA-256, real CLI startup, lease and host integration tests; production shared-broker authentication passed
Search and Matrix image / audio / video tools Matrix MCP and tool assembly restored; search is independent of mcode-tools MCP configuration and auth-isolation tests; actual search passed; generation requests not run
Official and local plugins; runtime GitHub importer /plugins and mcode plugin manage official installations and discovered local packages; arbitrary marketplace registration and GitHub URL import are not exposed in the CLI/TUI. See plugin management Plugin application, source-page, and offline local CLI tests; actual official catalog read passed
Managed connectors Cloud client, permissions, and invocation adapters restored Connector runtime and cloud transport tests; actual tool discovery passed; business writes not run
Updates Update entry points, install-source detection, and signature verification restored; public packages use public npm Update application and service tests; no real installation / upgrade on the development machine
Feedback and diagnostic uploads Reviewed feedback text and minimized diagnostic summaries Synthetic fixtures exercise session collection and capture/decode the final ZIP upload; no live uploads or real content
Automatic LLM error reports Disabled by default; requires telemetry.diagnostics opt-in; bounded diagnostic facts minimized before encryption Synthetic provider errors and native Headers; final HTTP batches captured and decrypted locally
Telemetry Usage, metrics, and diagnostics each disabled by default with separate opt-ins; MCODE_DISABLE_TELEMETRY / DO_NOT_TRACK override all channels; usage-event identifiers removed Privacy regression tests intercept and decode requests locally; no live upload
Auto permissions Cloud classifier restored; local rules and confirmation on failure retained Classifier client, permission facade, and sandbox tests
Model catalog Online catalog and bundled snapshot fallback restored Actual build boundary checks and offline startup; online catalog contents not accepted
Files, shell, subagents, sessions, headless, ACP Actual runtime retained BYOK, file reads, session resume, ACP, sandbox, and status protocol tests
Built-in skills, MCP, plugin tools Original TUI assets and activation conditions retained Asset build, plugin, and MCP tests; no claim that every skill has passed a real task

Local Bash execution

When the current turn includes native task_output, foreground Bash waits up to 60 seconds before returning the same command's background task ID. Its total command timeout defaults to 600 seconds and is capped at 600 seconds; a shorter requested timeout applies. Backgrounding and output reads preserve the original deadline. Without native task_output, Bash stays in the foreground with a 120-second default and a 300-second cap, and its schema omits run_in_background. Explicit background commands use the requested timeout; when omitted, the existing 30-minute runtime watchdog applies.

Only exit code zero is success. Results retain available exit, signal, timeout, cancellation, and partial-output facts. Large output keeps its original beginning and end within a 24 KiB first-response text budget, with a full-log reference when persistence succeeds. task_output reads use byte offsets; a successful read can report a failed command. Stop failures and incomplete logs are reported separately. An optional description supplies the TUI summary while execution and permission checks continue to use the original command.

Stopping a conversation and its background work

An explicit user stop (Esc in the TUI) also cancels background Bash commands and subagents owned by that conversation in the current process, including activated task_append continuations and work spawned by their child turns. Completion notices for those tasks no longer automatically wake the conversation. Their terminal results remain unread so the next turn can inspect them through the background-task reminder and task_output.

Leaving a conversation with /clear or a session switch still stops its current turn and pauses its Goal and queued instructions, but leaves background work running. This pause also applies when unconsumed steering must fall back to the queue after a history-write failure. A stop rejected because it names an older turn does not cancel the current turn's background work.

Only locally owned running tasks are canceled; tasks held by another runtime process are left alone. Child cleanup follows the task's owning child turn, so a later continuation in the same child session is preserved. Delivery suppression is held in process memory. If a turn does not release within the abort timeout, work created later during its shutdown is outside the stop boundary. A delivery already being admitted can also race with a stop.

Offline regression tests cover cascade ownership, delivery suppression, late task creation, repeated stops, append completion/shutdown, and queue fallback. The host integration tests use a scripted runner and real background processes; they do not establish live-model or cross-platform acceptance.

Skill directory links

Workspace .agents/skills, .claude/skills, and .minimax/skills support directory symlinks, both for the entire skill root and for individual skill directories. Targets may live outside the workspace. Existing external-source enable settings and duplicate-name priority still apply. Linked directories are watched for SKILL.md creation and edits; broken links are skipped. SKILL.md itself must remain a regular file.

Slash-command parity (TUI ↔ Web UI)

The webui talks to the same engine, but it only handles a strict subset of TUI slash commands itself (its LIB/local commands sourced from the mcode engine pass-through). MCODE_ACP_CAPABILITIES = {set_mode, set_config_option, cancel, activate, fork, resume, delete=false, load, close, list, new, prompt} (packages/webui/server/lib/mcode-rpc.js).

TUI built-in slash commands (packages/tui/src/application/command-descriptors.ts)

These appear in TUI's /help and are sent as prompt content over acp:

  • /help — Show available commands
  • /new — Start a fresh session in the current workspace
  • /model [provider/model[#variant]] — Choose a model
  • /status — Show account and model status
  • /doctor — Check the local config file
  • /context — Show the Runtime-owned context snapshot
  • /skills [filter] — List built-in and user Skills
  • /mcp [filter] — Inspect MCP capabilities and project configuration
  • /usage — Show session usage
  • /compact [instructions] — Shorten the active conversation
  • /export — Export the current Session as Markdown

WebUI-local commands (packages/webui/server/lib/interaction/commands.js#LOCAL_HELP_FALLBACK)

The webui ships its own local slash handler in server/lib/interaction/commands.js#LOCAL_HELP_FALLBACK and a local /help fallback; commands flow through server/lib/slash.js#handleLocalSlash → handleCmdCommand:

  • /new, /clear, /status, /sessions, /review, /help, /usage, /stop, and the goal commands /goal, /goal-done, /goal-blocked.
  • /review is the TUI parity path — it reads the current git status via server/lib/git.js#gitStatus and prints a staged / unstaged / untracked overview.
  • /clear and /new are destructive: they go through the authorize("slash.clear") gate (5-minute default timeout, fail-closed) before the underlying mutation runs.
  • The full catalogue of mcode-sourced pass-through commands is fetched via acp on session connect; the webui flattens {mcode:[…], extra:[…]} into the slash palette (webapp/test/slash-commands.test.ts).

Side-by-side command parity

Status legend: ✅ wired · ⚠ partial / path differs · ❌ no path · 🚧 requires engine-side support · — n/a.

Command TUI Web UI Notes
/help ✅ ✅ Web UI: bodyHelp lists webui-local + acp pass-through; the webui-local half is command-registry.js#CMD_BUTTON_COMMANDS on both the live and the fallback path, so it matches what POST /api/cmd accepts. LOCAL_HELP_FALLBACK covers the window before ensureMcodeCommands() has returned.
/new ✅ ✅ Web UI: B03 authorize gate (slash.clear); clears chat and rebinds the session.
/clear ✅ ✅ Web UI: B03 authorize gate; clears chat without touching mcode.
/status ✅ ✅ Web UI: bodyStatus prints account / model / quota via runUsageQuery.
/sessions ⚠ ✅ Web UI exposes a sidebar session tree (webapp/components/session-tree.tsx) backed by GET /api/sessions. TUI's /sessions resumes the chosen session.
/model ✅ ⚠ Web UI has a model selector in the toolbar (composer.tsx data-testid="model-selector-trigger") that posts to /api/set-model; no dedicated /model slash command in the webui local handler.
/doctor ✅ 🚧 Reads the local config file; webui has no equivalent routine.
/context ✅ ⚠ Web UI shows the context meter (components/context-meter.tsx) live; no slash command — the state.context snapshot is the only API.
/skills ✅ ⚠ Skill autocomplete shows up in acp pass-through via flattenAvailableCommands; no webui /skills slash body.
/mcp ✅ 🚧 Web UI has no equivalent.
/usage ✅ ✅ Web UI: runUsageQuery (server/lib/usage.js); also surfaces a forecast at GET /api/usage/forecast.
/compact ✅ 🚧 Web UI has no equivalent.
/export ✅ ✅ Web UI: GET /api/sessions/:id/export?format=md|json[&download=true] (B03 authorize-gated, written-ahead audit).
/stop ⚠ ✅ Web UI: dedicated POST /api/stop (gentle acp cancel + SIGTERM/2 s SIGKILL fallback); also exposed via toolbar Stop button.
/review ✅ ✅ Web UI: bodyReview → gitStatus → emits a staged/unstaged/untracked overview into chat. TUI parity.
/goal and /goal-* family 🚧 ✅ Webui-only (cs.goal = {active, text, status, duration, startTs}); sets a session goal and forwards to mcode via rewriteContent.

The webui does not yet expose /fork, /resume, or a "rewind last turn" action — those acp methods (fork, resume) are reported by MCODE_ACP_CAPABILITIES but no webui route wraps them. (Since M3-B4 that table is no longer what GET /api/protocol/capabilities returns; the endpoint serves the engine's declared 14-key capability object instead. The table is still exported and still pinned by test/lib/mcode-rpc.check.mjs.)

ACP Skill commands

ACP clients receive enabled Skills alongside built-in slash commands when a session is created, loaded, resumed, or forked. Discovery uses the session's Agent and workspace, including installed plugin Skills. Command names come from the runtime Skill roster (for example, /review or /plugin:review), not the installation package name. Select a command and append instructions to invoke it through the normal Agent turn. /skills [filter] lists the session's available Skills.

Built-in command names take priority over conflicting Skill names. Disabled, duplicate, and invalid command names are omitted. If Skill discovery fails, built-in commands remain available. Reopen the session after installing or enabling Skills to refresh its command menu. Protocol tests cover command discovery and prompt forwarding; this does not establish live Zed or model acceptance.

The webui tests the flattenAvailableCommands / filterCommands pipeline against the same engine-supplied shape (webapp/test/slash-commands.test.ts); the palette is capped at 8 entries by default.

Engine capability surface (/api/protocol/capabilities)

GET /api/protocol/capabilities returns:

{
  ok: true,
  mcodeVersion,
  mcodeName?,
  mcodeTitle?,
  // The engine's DECLARED 14-key capability object, served by
  // packages/webui/server/engine/capability-reads.js. Before M3-B4 this
  // field carried MCODE_ACP_CAPABILITIES (the ACP wire table in
  // packages/webui/server/lib/mcode-rpc.js), which is still exported
  // there and still a true statement about the ENGINE's ACP surface.
  capabilities: <14-key declaration>,
  capabilitiesProvider,        // which provider's declaration answered
  capabilitiesProviderFor,     // "transport" | "default" (see API.md)
  capabilitiesUnavailable,     // the degradation roll-up
  notes: {
    set_mode, set_config_option, cancel, activate, fork,
    load, list, close, new, prompt,
    session.delete → false (no engine handler; uses SQL on local_runtime_*)
  }
}

The webui uses the set_mode and set_config_option capabilities through the toolbar's permission-mode selector; cancel is what POST /api/stop sends first (with SIGTERM/SIGKILL only as a fallback when the notification cannot be delivered).

Desktop boundary

Background workspace indexing is removed from this distribution. Runtime startup and conversation turns do not collect workspace snapshots, create workspace ZIP archives, or upload/retry them for cloud indexing. The semantic workspace search tool and its enablement policy are also removed; a saved indexing preference cannot reactivate them. Existing indexing records are left inert. User-directed file reading, search, and Git operations remain available.

Desktop features not enabled by default in the original TUI are outside this restoration commitment: remote encrypted prompt updates, cloud session handoff, and native Electron desktop control. Earlier descriptions incorrectly counted removal of those sources as TUI feature loss. Desktop HTTP services and internal generated IDL remain outside the public repository.

Website deployment tools and routes follow the original tool set. Successful deployment still depends on login, service permissions, and backend support; no real deployment has been performed.

Release boundary

Production client paths are restored without internal registries or test-service addresses as build dependencies. Real account behavior, quota, billing, model quality, media generation, and deployment require separate acceptance. Mocks, protocol fixtures, successful builds, and the source import do not replace that evidence. Public npm latest was 0.4.12 on 2026-09-18; that metadata observation does not validate the installed package or deploy a backend. See verification records and the historical release audit for recorded acceptance results.

Diagnostic upload privacy

Automatic LLM error reports use an allowlist before buffering and encryption. Schema 3 retains a numeric release version, the fixed metric error category, known error names/codes, HTTP statuses (100–599), and bounded error/cause/properties relationships. It excludes error messages, stacks, headers (including native Headers), prompts, request/model metadata, URLs, arbitrary property names/values, and binary data. Unknown event types are dropped; the event type and code location are fixed at the reporter boundary. Encryption remains a transport layer, not redaction. Authenticated transport still uses the signed-in account token and user ID.

The separate TUI incident reporter uses schema 2. It retains a fixed event category, phase/severity/impact, handled flag, generated IDs, timestamps, numeric release/Node versions, known platform/architecture, known error names/codes and HTTP status, and up to 20 fixed breadcrumb names with an optional phase. It excludes free-form error text, stack/cause/aggregate contents, caller context, component/operation labels, terminal/OS strings, and original code locations. Fingerprints are derived only from the minimized category and error facts, so grouping is intentionally coarser. Legacy schema-1 pending and sent files are deleted best-effort at startup and never uploaded; current-schema pending files are projected again before encryption and replaced with their minimized copy before being marked sent. The gateway wire encryption format is unchanged. Production ingestion of schema 2 has not been tested.

Feedback review describes the actual upload: user-reviewed feedback text, bounded client/platform and optional session metadata, and diagnostic counts, not original session files. Recognized credentials are redacted from the reviewed description; arbitrary personal text in that description is still sent, so review it before submitting.

Every collected diagnostic artifact, including prioritized session artifacts, is projected through diagnostic-counts-v1. The ZIP contains counts of known roles, states, error types/codes, log levels and HTTP statuses, plus parsing/omission indicators. JSON/JSONL can contribute these facts; malformed data, binary attachments and free-text logs contribute no raw content. Source filenames and paths are replaced with opaque archive names; only a small fixed list of artifact kinds (such as messages.jsonl) is retained. Unrelated session manifests are not collected. Local source files are unchanged.

There is no raw-attachment upload option in this flow. Prompts, conversation text, tool arguments/results, command output, workspace excerpts, raw errors and unknown fields are excluded even if they contain no recognizable credential pattern. This intentionally reduces diagnostic detail: reproducing an exact response or inspecting an original stack is not possible from these uploads. Future raw attachments would require a separate, explicit review and consent surface describing their contents and scope.

The regression tests use temporary synthetic files and intercepted HTTP only. They decrypt the final automatic-report request as a receiver would, and unzip the actual feedback PUT body after real session-report collection. They do not validate production ingestion of the new schemas, retention policies, live services, other telemetry paths or other platforms.

Interactive startup model

Use mcode -m <provider-id>/<model-id> or mcode "Fix the failing tests" --model <provider-id>/<model-id> to select the model for the startup Session. References use the same syntax as exec --model, including custom_provider:<id>/<model-id>, model IDs containing /, and the optional #variant suffix. The provider must already be configured and the model must be available to Runtime; invalid references fail before the initial prompt runs.

Without a resume option, this creates a new Session even when no prompt is given. With --session <id> or --continue, it updates the opened Session's model before submitting the prompt. The selection is saved with that Session, so later turns and resumes keep it. It does not change the global default, and /new returns to the configured default. Omitting --model preserves existing startup behavior. The untargeted --session picker cannot be combined with --model; provide an ID or use --continue instead. If opening or continuing a Session fails, the override and initial prompt are not applied.

Feature panels and chat restoration

In regular mode, independent feature panels occupy the complete visible terminal area, including short Rewind previews and scope pickers. Closing a panel restores the current conversation. Closing, replacing or shrinking a transient region restores the exposed chat rows. This includes inline selectors such as /theme, completion menus, multi-line drafts, image previews, queued messages, task and Goal summaries, welcome notices and status rows. Short documents refresh in place; history is reconstructed only when the smaller layout needs to bring scrolled rows back into view. This rule follows the rendered layout, including asynchronous updates, rather than requiring each close handler to request a special redraw. When background running content shrinks entirely within the current screen and the transient layout stays unchanged, the renderer keeps native scrollback and the Composer position stable. Freed rows temporarily remain blank at the top of the active screen and subsequent output reuses them. This avoids resetting the host's scroll position when a turn finishes. Redundant resize notifications with unchanged dimensions do not rebuild history. Viewport-only redraws erase rows in place so terminals that save a cleared screen to scrollback, including Apple Terminal, do not retain the old Composer, status line or duplicate transcript rows.

When a change removes or replaces text already in scrollback, the renderer still reconstructs the current session to avoid stale or duplicate history. Real resizes and image layout changes also retain the existing reconstruction behavior. A reconstruction clears earlier shell scrollback and can reset the host's scroll position; ordinary updates keep native scrolling and selection behavior.

Rewind and Fork history-loading hints disappear as soon as their lists are ready. Returning from a cancelled operation must not leave a stale loading message in the Composer. Rewind displays its completed result after a successful operation.

Temporary side conversations

Use /btw [question] (or /side [question]) to open a temporary side conversation while the main task continues. The side conversation inherits the latest complete prefix of persisted history. Completed tool calls retain all their results; an unfinished group of tool calls is excluded together. The selected boundary is fixed before creation, so later main-task output does not change that fork. Inherited tool calls are context and are not executed again.

Press Ctrl+/ to switch between the main and side views, or /parent to return to the main view. Press Ctrl+C on an empty Composer to discard the side conversation. Side conversations retain the main session's permission mode and remain hidden from /sessions and /resume.

Creation and activation failures record a bounded, redacted cause chain in the local session.side.failed diagnostic event. Feedback uploads still apply the existing diagnostic-counts projection; raw error text, stacks and session IDs are not added to the uploaded ZIP. Offline tests cover persisted tool histories, archives, concurrent parent output, side-session cleanup and local diagnostics; this does not establish native-terminal or live-model acceptance.

Select a plugin for a message

Type @ in the Composer to search files and installed, enabled plugins. Plugin candidates show their source so packages with the same display name can be selected independently. Choose a plugin with Tab or Enter, then describe the task. The Composer shows @Name and retains the plugin identity through editing, undo, prompt history, queued-message recovery, saved drafts, and /edit after a message is sent. Ctrl+C clearing/restoration and external-editor edits retain unchanged plugin bindings. If external edits make duplicate labels ambiguous, reselect those plugins in the Composer. Displayed messages remain readable; Runtime retains the original input separately when needed to recover the plugin identity for editing.

Selection applies to that message. Runtime checks the plugin's effective Skills, MCP tools, and App tools again for the turn and asks the Agent to prefer relevant capabilities. Selecting a plugin does not install or enable it. An unavailable selection is reported to the Agent rather than redirected to a same-named package.

Exec and ACP text prompts can use the durable linked form, for example [@Notes](plugin://notes%40local) summarize these files. The ID is the package name plus its local or official source; display labels do not determine the selection. Legacy whitespace-delimited @package-name text remains supported when it identifies exactly one effective plugin.