Skip to content

0.4.5: constant system prompt per session, no more full re-reads every turn - #83

Merged
mjwsolo merged 1 commit into
mainfrom
release/0.4.5
Sep 26, 2026
Merged

mjwsolo merged 1 commit into
mainfrom
release/0.4.5

Conversation

@mjwsolo

@mjwsolo mjwsolo commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

llama-server reuses its KV cache only for an identical prompt prefix. The system prompt was changing between steps, so the server re-read the entire conversation twice per turn (15-23 s at 30k tokens, measured in a real session):

  • fork session/system.ts: workspaceActive was recomputed per turn, so step 1 of a turn lacked the workspace block and step 2 had it. Now sticky per session.
  • plugin: the planning rule followed the same switch; it is now always present for the build agent. The open-todo list moves from the system prompt to the user turn.
  • fork provider.ts: short side requests (title generation) carry id_slot = last slot so they never truncate the conversation's cache on multi-slot servers (no-op on one slot).
  • fork tui: the code-intelligence panel says "Not running · /lsp to set up" instead of promising an automatic start that cannot happen without a download.
  • Bundled UI binary rebuilt from fork f65ef33d.

Measured on Qwen3.8 27B, three turns with tool calls: after the first tool call every request re-reads 23-310 tokens (0.3-0.8 s) instead of 6-7k (10-15 s).

tests/ui/prompt-scope.test.ts now states the new contract. The other 4 bun failures are the known pre-existing ones.

…ry turn (fork f65ef33d)

llama-server reuses its cache only for an identical prompt prefix, and the system
prompt changed between steps: the fork added the workspace block after the first
tool call of a turn and dropped it on the next, the plugin's planning rule followed
the same switch, and the open-todo list was rewritten into the system prompt on
every update. Each change re-read the whole conversation, twice per turn (15-23 s
at 30k tokens). Now: workspaceActive is sticky per session (fork), the planning rule
is always present for the build agent, todos ride on the user turn. Measured on
Qwen3.8 27B: 23-310 tokens (0.3-0.8 s) per request after the first tool call
instead of 6-7k (10-15 s). Also: short side requests pinned to the last server
slot; code-intelligence panel says 'Not running · /lsp to set up'.
Bundled UI binary rebuilt from fork f65ef33d (sha256 67582dda…).
@mjwsolo
mjwsolo merged commit c18d851 into main Sep 26, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant