docs(adr): ADR 0045 — think+format follows upstream's single pass by default; production runs two-pass - #392
Conversation
…default; production runs two-pass Open item 3 of the v0.34.4 fold record, written on the maintainer's word after the three hosts' data: - The code default stays upstream's single pass; OLLAMA_FORMAT_TWO_PASS=1 selects ADR 0004's flow. Production runs the switch (the maintainer's decision 2026-09-27; CUDA and gfx1151 deployed 2026-09-28). - The reason is MLX drafting: draftingEnabled is Grammar == nil || draftUnderGrammar, and production keeps OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0 (ADR 0033) for the qwen3.5-family retention. A single pass carries its grammar from token 0, so it never drafts; two-pass's pass one drafts. Measured: P0 thinks 1.46-1.7x faster than F0 on CUDA; on Metal the undrafted single pass leaves gemma4:26b 4 and 31b 2 cases unfinished against 1 and 0 for two-pass and for the single pass with the knob at 1; on MLX-CUDA and on GGUF the flows tie. - Records the two flows' budget semantics (num_predict, tools, raw generate, EOS in the thinking, MLX drafting, prefills, metrics), and proposes the switch's retirement condition: a single pass that drafts its thinking without the retention, gated by leak-repro5.sh, the drafting probe and Metal's think-on protocol. - ADR 0004 is marked superseded as the statement of which flow runs (it remains the two-pass reference); the retirement register row, the fold record's item 3 and the README's think+format row point at ADR 0045. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
gfx1151 review of ADR 0045. The gfx1151 facts are right, except one line that generalises qwen3.6.
|
…ploy checks both variables From the gfx1151 review on #392: - "Quality leans toward two-pass" generalised qwen3.6's counts. Across the five GGUF models there is no net direction: qwen3.6 leans toward two-pass, nemotron3 toward the single pass on the bbox contract (23/40 against 15/40, p ~ 0.1, not established), and the other three do not move. Checked against the fold record's gfx1151 think-on section. - Decision 3: gfx1151's deploy script also sets and checks both variables, and its compose file carries them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks, both applied in
|
|
The maintainer merged #392 into
|
This closes open item 3 of the v0.34.4 fold record, on the maintainer's word.
ADR 0045 records the two decisions the maintainer has made:
OLLAMA_FORMAT_TWO_PASS=1selects ADR 0004's flow.It rests on one interaction. On MLX,
draftingEnabledisrequest.Grammar == nil || draftUnderGrammar. Production keepsOLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0, because drafting under a grammar retains memory on the qwen3.5 family (ADR 0033). A single pass carries its grammar from token 0, so it never drafts. Two-pass's first pass has no grammar, so it drafts the thinking.The data, from the fold record and #375:
It also:
num_predict, tools, raw generate, EOS inside the thinking, MLX drafting, prefills, metrics), which item 3 asked for;leak-repro5.sh, the drafting probe and Metal's think-on protocol.Other files:
This covers the CUDA and gfx1151 deploys; the Metal host's deploy is decided separately.
ai-server/mlx-cuda🤖 Generated with Claude Code