Skip to content

context: compaction budget measured on this machine; per-slot ctx from the server (0.4.9) - #90

Merged
mjwsolo merged 7 commits into
mainfrom
feat/dynamic-context
Sep 27, 2026
Merged

mjwsolo merged 7 commits into
mainfrom
feat/dynamic-context

Conversation

@mjwsolo

@mjwsolo mjwsolo commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Two changes, both dynamic on the machine and model, no token literals.

  1. Compaction point from measured speed. The supervisor reads llama-server's own print_timing lines from server.log as the session runs (prompt length, prefill t/s, decode t/s per request) and publishes a budget in /status: the largest prompt length at which a full re-read finishes within 60 s and decode keeps at least half of its short-prompt speed, clamped to [n_ctx/4, n_ctx - one step's headroom]. Before measurements the warm-up replay's prefill speed is the prior; without that, the cap (today's behaviour). The fork provider sets it as limit.input, which overflow.ts already prefers. Evidence: on the owner's logs decode halves and prefill drops 4-6x between 16k and 64k prompts, and sessions were reaching 120k before compacting.

  2. One source of truth for the context window. After load the supervisor reads the per-slot n_ctx from /props (with -np N the flag is split N ways) and reports it as ctx; the launcher's fixed 32768 fallback is gone. This is the class behind the 'request exceeds the available context size' failures on 24 Sep.

Tests: 11 new unit tests (parsing, rules, prior, clamps, launcher config, supervisor status with a fake /props and a growing log); launcher and warm-up suites pass. Binary rebuilt at fork 86d2d971.

Live verification (isolated eval server, Qwen3.6-35B-A3B Q8, ten-turn task):

  • 32k context: ctx read from /props = 32768, budget published at 16384 (measured basis, 13 samples), runtime received it; no compaction needed (max prompt 6.7k).
  • Default context: budget moved 97280 -> 73728 as deep samples arrived.
  • Forced (16k context, re-read cap 1 s): budget 8192 (floor = 2 x prefix 5919, bounded by the cap); exactly one compaction when the prompt crossed 8192 (max 8433); zero 'exceeds the available context' errors; all ten turns completed.
  • Two defects found and fixed by that forced run before shipping: a floor below the fixed prompt prefix compacted on every step; the prefix must come from the first main request, not the smallest prefill.
  • Control: Qwen3.8-27B on the same branch reuses the prefix perfectly (24-29 tokens per turn) and passes. Qwen3.6's turn-two full re-read reproduces on main and is model-specific (checkpoint-based reuse), tracked separately.
    Checks 10/10 with the new supervisor and binary.

…m the server

The supervisor reads the per-slot context from llama-server's /props after
load, parses the server's own print_timing lines from server.log as the
session runs, and publishes in /status a budget: the largest prompt length at
which a full re-read finishes within REREAD_MAX_S and decode keeps at least
DECODE_FLOOR of its short-prompt speed, clamped to [n_ctx/4, n_ctx - one
step's headroom]. Before measurements the warm-up replay's prefill speed is
the prior; without that, the cap. The launcher sets compaction.reserved to 0
(the budget already holds the headroom) and derives its last-resort context
from the config knob instead of a literal 32768.
Forcing the budget to n_ctx/4 at a 16k context put it below the system prompt
and tool schemas (about 4.8k tokens), so no compaction could bring the prompt
under budget and the runtime compacted on every step. The floor is now the
larger of n_ctx/4 and twice the fixed prefix (the smallest cold prefill seen),
bounded by the cap.
…efills

The smallest large prefill is wrong once a compaction has happened (those
prefills are the summary plus recent turns, well under the system prompt),
and a 16k run still compacted on every step. The first main request of a
session is among the first few large prefills after load; side requests are
small and the warm-up replay equals the prefix, so their maximum is it.
@mjwsolo
mjwsolo merged commit 9ff421f into main Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant