context: compaction budget measured on this machine; per-slot ctx from the server (0.4.9) - #90
Merged
Merged
Conversation
…m the server The supervisor reads the per-slot context from llama-server's /props after load, parses the server's own print_timing lines from server.log as the session runs, and publishes in /status a budget: the largest prompt length at which a full re-read finishes within REREAD_MAX_S and decode keeps at least DECODE_FLOOR of its short-prompt speed, clamped to [n_ctx/4, n_ctx - one step's headroom]. Before measurements the warm-up replay's prefill speed is the prior; without that, the cap. The launcher sets compaction.reserved to 0 (the budget already holds the headroom) and derives its last-resort context from the config knob instead of a literal 32768.
…e measured budget)
Forcing the budget to n_ctx/4 at a 16k context put it below the system prompt and tool schemas (about 4.8k tokens), so no compaction could bring the prompt under budget and the runtime compacted on every step. The floor is now the larger of n_ctx/4 and twice the fixed prefix (the smallest cold prefill seen), bounded by the cap.
…efills The smallest large prefill is wrong once a compaction has happened (those prefills are the summary plus recent turns, well under the system prompt), and a 16k run still compacted on every step. The first main request of a session is among the first few large prefills after load; side requests are small and the warm-up replay equals the prefix, so their maximum is it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two changes, both dynamic on the machine and model, no token literals.
Compaction point from measured speed. The supervisor reads llama-server's own print_timing lines from server.log as the session runs (prompt length, prefill t/s, decode t/s per request) and publishes a budget in /status: the largest prompt length at which a full re-read finishes within 60 s and decode keeps at least half of its short-prompt speed, clamped to [n_ctx/4, n_ctx - one step's headroom]. Before measurements the warm-up replay's prefill speed is the prior; without that, the cap (today's behaviour). The fork provider sets it as limit.input, which overflow.ts already prefers. Evidence: on the owner's logs decode halves and prefill drops 4-6x between 16k and 64k prompts, and sessions were reaching 120k before compacting.
One source of truth for the context window. After load the supervisor reads the per-slot n_ctx from /props (with -np N the flag is split N ways) and reports it as ctx; the launcher's fixed 32768 fallback is gone. This is the class behind the 'request exceeds the available context size' failures on 24 Sep.
Tests: 11 new unit tests (parsing, rules, prior, clamps, launcher config, supervisor status with a fake /props and a growing log); launcher and warm-up suites pass. Binary rebuilt at fork 86d2d971.
Live verification (isolated eval server, Qwen3.6-35B-A3B Q8, ten-turn task):
Checks 10/10 with the new supervisor and binary.