0.4.5: constant system prompt per session, no more full re-reads every turn - #83
Merged
Merged
Conversation
…ry turn (fork f65ef33d) llama-server reuses its cache only for an identical prompt prefix, and the system prompt changed between steps: the fork added the workspace block after the first tool call of a turn and dropped it on the next, the plugin's planning rule followed the same switch, and the open-todo list was rewritten into the system prompt on every update. Each change re-read the whole conversation, twice per turn (15-23 s at 30k tokens). Now: workspaceActive is sticky per session (fork), the planning rule is always present for the build agent, todos ride on the user turn. Measured on Qwen3.8 27B: 23-310 tokens (0.3-0.8 s) per request after the first tool call instead of 6-7k (10-15 s). Also: short side requests pinned to the last server slot; code-intelligence panel says 'Not running · /lsp to set up'. Bundled UI binary rebuilt from fork f65ef33d (sha256 67582dda…).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
llama-server reuses its KV cache only for an identical prompt prefix. The system prompt was changing between steps, so the server re-read the entire conversation twice per turn (15-23 s at 30k tokens, measured in a real session):
session/system.ts:workspaceActivewas recomputed per turn, so step 1 of a turn lacked the workspace block and step 2 had it. Now sticky per session.provider.ts: short side requests (title generation) carryid_slot= last slot so they never truncate the conversation's cache on multi-slot servers (no-op on one slot).Measured on Qwen3.8 27B, three turns with tool calls: after the first tool call every request re-reads 23-310 tokens (0.3-0.8 s) instead of 6-7k (10-15 s).
tests/ui/prompt-scope.test.tsnow states the new contract. The other 4 bun failures are the known pre-existing ones.