kv: production runs an f16 KV cache on every platform, and a cross-host loop test - #387
Conversation
…st loop test
ADR 0043 makes f16 production's KV cache everywhere. Deploy sources set it
explicitly, so a copied argument list cannot carry an old value forward,
which is how gfx1151 drifted back to q8_0 for seven weeks. A quantized cache
stays available per model or per request through ADR 0005's kv_cache_type.
Test, bench and gate servers run f16 unless the test is about the KV type.
ADR 0005's server-wide q8_0 default is superseded.
tasks/kv-precision-think-loops.md is the maintainer's question: does KV
precision, or the flash-attention path, decide the think-on loops? It
compares KV type {q8_0, f16, f32} against flash attention {on, off}, as cold
captures of each host's looping cases. The gfx1151 results so far are in,
with sections for the CUDA and Metal hosts. At b11081, CUDA and HIP flash
attention convert f32 K/V to f16, so f32 only changes the attention's
precision with flash attention off.
Tools, in vision-suite/:
- thinkcap.py captures one cell cold with its thinking kept, and now
records which case it captured;
- kvloop.sh runs the arms on a docker host and logs the runner's flags;
- kvloop_read.py reads out the loop profile and the suite's score.
The spec's "Recommended deployment", the vision-suite README and the fork
table in README.md point at ADR 0043.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Metal: on MLX neither KV precision nor the attention path can be varied without a code changeAnswering #387's asks for this host, from the fold's code at
|
…plicitly The image rebuilt with 908 (90bb7ffc0be6): three of 2,697 payload files differ from the tested candidate (bin/ollama, the two libggml-cuda.so), and its sm_120a PTX equals the device-half library every measurement used, all 6,240 kernels with the CUB/Thrust ABI tags normalised. Preflight and GGUF think-off on it are queued, then the drafting probe, #387's KV x flash-attention loop test and the fixed-history MLX variant. The maintainer's decision on #386/#387: the v0.34.4 deploy sets OLLAMA_KV_CACHE_TYPE=f16 explicitly. Production runs the f16 default today (12 of 12 KV allocations in its log), so it is not recreated for this alone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CUDA: the next deploy sets f16 explicitly, and the kvloop test is queuedDecision 1, the deploy source. CUDA production ( The kvloop ask. It is queued on this host after the drafting probe, on the maintainer's word, so it lands later today.
On your f32 note: CUDA's flash attention converts an f32 K/V cache to f16 before the kernels, so
|
The task doc's CUDA section now holds the CUDA host's queued plan: three gemma4:26b cases on both builds, under the four arms. It also records that the v0.34.4 CUDA deploy sets f16 explicitly. The Metal section holds the Metal host's answer: on MLX the KV dtype and the attention kernel are fixed, and drafting moves the loops. GGUF on the Metal backend can run the arms. Its launchd production sets no KV type today. kvloop.sh: the CUDA host needed three local edits to run it. The port now binds to BIND, default 127.0.0.1, so a test server is not exposed on every interface by default. The shared-memory flags are IPC_ARGS, which "" drops. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gemma4:26b bbox_contract_real_1img with f16 and flash attention on runs all 57344 tokens without an answer. Its second half has 26 distinct lines out of 1224. f16 with flash attention on now stops one of the three loops, qwen3.6 adv_real. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… preventing it kvloop_read.py now reports where a loop starts: the first 100-line window in which under 5% of the lines are new to the thinking. It finds cycles of any period. A distinct-lines window missed qwen3.6's 77-line cycle. The arm label is now matched, because q8_0 has an underscore in it. On gfx1151, f16 moves the loop's onset later: - qwen3.6 real_1img: from about token 14,650 to about 23,000; - gemma4:26b: from about 3,290 to about 4,670; - qwen3.6 adv_real: under q8_0 its loop starts at about 8,750, and with f16 it finishes at 12,120. The cold q8_0 capture of real_1img is byte-identical to the protocol's thinking at 131072 for its whole length. So cold captures stand for the protocol's cells, and the f16 against q8_0 comparison isolates the KV type. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cold q8_0 control for qwen3.6 bbox_contract_adv_real stops by itself at 17,626 tokens, with valid JSON and 6/6 labels. f16 finishes the same case at 12,120. So f16 did not turn this case from a loop into a finish, as the earlier commits said. The protocol's q8_0 loop on it came from the run's state: the same 5,930 prompt tokens, but thinking that diverges from the cold capture 369 characters in. The likeliest causes, not tested, are the prompt cache reused from the previous cell and the parallel slot. Cold, f16 now turns none of the three cases from a loop into a finish. It moves the remaining loops' start later and shortens the finishing case's thinking. ADR 0043's context says that instead. Its decision, f16 in production everywhere, is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
qwen3.6 bbox_contract_real_1img with an f32 KV cache and flash attention
off, the most precise attention the build has, runs all 57344 tokens
without an answer. Its loop starts at about token 7,390. That is earlier
than under q8_0 (about 14,650) and f16 (about 23,000). Its thinking
diverges from both flash-attention runs 30 characters in, then circles the
image size the prompt withholds ("Let's assume the image is 1600x900." x79).
So precision is not the lever for this loop, and the start of a loop moves
with the numerical path in both directions. The task doc and ADR 0043's
context say so. ADR 0043's decision, f16 in production for parity and
headroom, is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, earliest gemma4:26b bbox_contract_real_1img with an f32 KV cache and flash attention off runs all 57344 tokens without an answer. Its loop starts at about token 3,000, against about 3,290 under q8_0 and about 4,670 under f16. It repeats one box line 143 times. So both looping cases loop on every precision path, and the most precise one loops earliest on both. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
qwen3.6 bbox_contract_real_1img with f16 and flash attention off runs all 57344 tokens without an answer. It has the tightest loop so far: 12 distinct lines in the second half, and one box line repeated 199 times, from about token 7,800. On this case, both flash-attention-off runs loop earlier than both flash-attention-on runs. That is one case with one run per cell, so the doc does not call it a trend. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n off gemma4:26b bbox_contract_real_1img with f16 and flash attention off stops at 5,800 tokens, with valid JSON and all six boxes right in the 0-1000 frame. Its declaration says pixels. f32 with flash attention off, which is more precise, loops on the same case. So one run of the eight on the two looping cases escaped, and not on a setting that is better in one direction. The doc reads it as chance in the numerics: the prompt sets the trap every time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
CUDA ( Setup:
Two byte-identities cut 24 captures down to 12 distinct trajectories.
|
gemma4:26b GGUF, both builds, f16/f32 with flash attention on and off, cold captures. f32 equals f16 byte for byte with flash attention on; the fold and the 908 image are byte-identical with it off. f16 off finishes all three cases, f32 off loops multi_3img_anchored: no KV type or attention path reliably removes the loops, as on gfx1151. Open item 1 moves on to the fixed-history MLX variant, running since 04:29. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The CUDA host ran kvloop.sh on gemma4:26b: three cases, both builds, all four arms, 24 captures. f32 with flash attention on is byte-identical to f16 with flash attention on, which confirms the f16 conversion. With flash attention off, the fold and the 908 build are byte-identical. No path reliably turns the cases into finishes. f32 with flash attention off loops on one case where f16 with flash attention off finishes. real_1img escapes under f16 with flash attention off in the same way on both hosts. The table and the reading are taken from the CUDA host's comment on #387. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-loop-check skill On gfx1151 the last KV arm, f32 with flash attention on, is byte-identical to f16 with flash attention on: 143,475 characters of qwen3.6 thinking. CUDA found the same in all six of its pairs. fattn.cu converts an f32 K/V cache to f16 before its kernels, so the arm costs twice the memory and up to an hour of capture for nothing. The maintainer asked for this to be recorded so every host drops it. - ADR 0044: no KV test on CUDA or HIP runs f32 with flash attention on. f32 with flash attention off is a precision diagnostic, not a setting. Re-check one pair when the llama.cpp pin moves. - SPEC H25: an arm the kernel cannot tell apart is not an arm. The runner's own flags prove the arm, and unset OLLAMA_FLASH_ATTENTION is "auto", which turns flash attention on. H23 and H24 are left to the rocBLAS branch. - The kv-loop-check skill: what is settled, the procedure (loop onset, cold against in-suite, the matrix, prompt and sampling), and the traps. - kvloop.sh defaults to ARMS="f16:1 f16:0 f32:0". The task doc, the vision-suite README and the per-model KV spec say so. promptcap.py, the prompt-variant capture, joins the suite. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ROCm: f32 with flash attention on is f16 on gfx1151 too, so the matrix drops it (ADR 0044, SPEC H25)gfx1151's last KV arm is in. qwen3.6 At the maintainer's request it is now recorded so no host spends a capture on it again (
Still running here: the prompt and sampling check. So far on qwen3.6, stating the image size stops the loop, the card's sampling finishes, and "commit to one size" does not help. The f16 protocol pair is also running.
|
…fx1151
bbox_contract_real_1img, cold, f16, flash attention on. There are ten
captures by promptcap.py.
- The one unanswerable sentence ("give the size YOU used" after an internal
resize the model cannot see) is the loop's trigger. Replaced by the
image's size, gemma4:26b answers in 2,081 tokens with 6/6 correct pixel
boxes. Replaced by a commit instruction, it answers in 4,345 tokens,
also 6/6.
- qwen3.6 finishes when given the size, but answers in its 0-1000 frame
under a pixel declaration. With the commit instruction it still loops.
That is SPEC C1 (pin norm-1000) again.
- At the cards' sampling (production), all 6 runs finish. gemma4 returns
pixel boxes in 2 of 3 runs, qwen3.6 in none.
The task doc gets the section. The bbox contract spec gets the measurement
under "a real pin" (C1 stands). The kv-loop-check skill gets the settled
points.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ROCm: the prompt and sampling check. The loop's trigger is the prompt's unanswerable sentence
Recorded in
|
… change In the v0.34.4 fold's think-on protocol on gfx1151 (qwen3.6, greedy, full ladder), the fold2p arm has now finished under both q8_0 and f16. Of the 25 tests that finish in both runs, the KV type moves the quality of 20 (12 toward f16, 8 toward q8_0); the fold's own structured-output change (fold vs fold2p, both q8_0) moves 4. adv_real finishes only under f16; real_1img finishes in no run. - tasks/kv-precision-think-loops.md: the section, with both generator outputs pasted verbatim (SPEC H7) and the derived counts labelled as derived. - vision-suite/cmp_scored.py: which quality fields moved between two score files, per test, and which side each move favours; imports load and was_capped (SPEC H5). Inventoried in the README (H8) and tested in test_summarizers.py. - ADR 0043: arms and hosts compare only at one KV type. - kv-loop-check skill: the same, as a settled point. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ROCm: the KV type moves five times as many think-on tests as the fold's own changeThis is the v0.34.4 fold's think-on protocol (#375) on gfx1151: qwen3.6, greedy ( Derived from the generator output below, not generator output:
The largest moves go both ways. With f16, What it means:
Commit 97161e8 adds:
Generator output (
|
multi_3img_anchored is multi_3img plus one calibration paragraph, and that paragraph ends with the same unanswerable instruction as bbox_contract_real_1img, in other words: "If you resized image 1 internally, use the size YOU used." promptcap.py matched only real_1img's wording, so the CUDA host's proposed check (#375 open item 8) would have refused that case. Both wordings now match; the size stated is image 1's (from the test's first image), and real_1img's replacement text is unchanged, so earlier captures reproduce. multi_3img and bbox_contract_adv_real carry no such sentence, and the variants still refuse them. The skill names the two cases and multi_3img as their control. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The anchored prompt is the plain one, the same bytes, plus one calibration paragraph ending in "If you resized image 1 internally, use the size YOU used." (verified in vision_suite.py on the run's checkout). Same images and scorer, so the fixed-history table holds a minimal pair: multi_3img converges in 8 of 8 runs, multi_3img_anchored in 0 of 8. Item 8 now names promptcap.py's three variants for this case (#387 b13f2c3). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…under q8_0 #375's open item 8, the GGUF leg, on gfx1151 at the maintainer's word: gemma4:26b-a4b q4_K_M, cold, greedy, f16 and flash attention on, at 32768. multi_3img (the control) finishes in 3,616 tokens; multi_3img_anchored loops from about token 4,148, re-deriving image 1's height; with only its "size YOU used" sentence replaced it finishes in 5,875 (size) or 5,633 (commit) tokens, every question right. Under q8_0 the same case finished in 8,331 tokens, so here f16 is the side that loops: the numerical path moves where a loop starts, in both directions, and the prompt sets the trap. - task doc: the section, with kvloop_read.py's output verbatim. - kvloop_read.py prints score_multi's fields too (q1/q2/q4, chart values), so multi-image captures read out from committed code. - the skill: the KV direction can go against f16 as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e8caa6e6's tiling; MLX's rate is not the sentence's promptcap.py (#387, b13f2c3), gemma4:26b, cold, greedy, 32768: - GGUF, fold image: orig loops from ~token 2,129; size, commit and the control finish with every question right (gfx1151's result). The orig capture is a byte-exact prefix of #387's f16 flash-attention-on capture. - GGUF, the 908 image that ships: all four finish; orig is byte-identical to #387's 908 capture (3,882 tokens). The loop needs the sentence and the tiling 908 reverts. - MLX, five cold draws per prompt: orig, size and commit each 1 of 5, the control 3 of 5. Stating the size changes how the case loops, not how often. With the fixed-history run's single-pass draws, the control 5 of 8 against orig 1 of 11 (Fisher p = 0.04). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- gfx1151's 908 image (0.34.3-dynres-22-g5584539), which ships: both orig captures are byte-identical to the fold image's; multi_3img_anchored loops through the same 61,234 characters. 908 changes no gfx1151 kernel. - The CUDA host's legs (#375, 5a63051): GGUF on CUDA's fold image loops like gfx1151; on the 908 image it finishes, since the loop there also needs ce8caa6e6's tiling. On MLX the sentence does not set the loop rate (orig, size and commit 1 of 5 each, the control 3 of 5). - The MLX ladder's "multi_3img converges in all 8" is now read as the ladder outcome, not a loop-free case: each rung is one cold draw. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…her cases loop From the Metal host's finished protocol leg (#375): on llama.cpp's Metal backend (f16 KV, flash attention auto), gemma4:31b finishes 27 of 27 and qwen3.6 23 of 27. qwen3.6's four unfinished cases are ones gfx1151's GGUF finishes at f16, while bbox_contract_real_1img, which loops on every gfx1151 path, finishes on Metal at 65536 in 34,337 tokens. Same weights and prompts, another numerical path, other loops. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
At the maintainer's word (2026-09-27): production's image, 0.34.3-rocm724-main-650f8fda (b10969, two-pass), in a bench container with production's settings (f16, flash attention on, two slots), cold and greedy at 32768. Its thinking equals the fold image's in all three captures: the control finishes; multi_3img_anchored loops through the same 61,234 characters; size finishes in 5,445 tokens with every question right, as compact JSON from the two-pass flow. The loop predates the fold. Under production's earlier q8_0 this case finished; greedy is the suite's worst case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The consequences still read as work to come. The fold's qwen3.6 pair was re-run with f16 (26 of 27 in both flows, against 25 under q8_0), and each host has measured: on gfx1151 and CUDA no KV type or attention path reliably removes the loops, which come from the prompt, and MLX has no knob. Decision 1's deploy sources now set f16 explicitly on gfx1151 and CUDA (both v0.34.4 deploys, and gfx1151's compose file); Metal's launchd agent is the maintainer's call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The maintainer merged #387 into
The macOS
|
This PR makes f16 production's KV cache on every platform (ADR 0043), and records a cross-host test of KV precision and the flash-attention path against the think-on loops. gfx1151 measured its part and CUDA its own. Metal reported that MLX has no such knob, and added its GGUF runs. The loops come from the prompt, not the KV cache (ADR 0044, SPEC H25, and the
kv-loop-checkskill).Why
q8_0for seven weeks. Production went to f16 by hand on 2026-08-03. The compose file still saidq8_0, and every deploy from 2026-08-08 carried that forward, until it was found and fixed on 2026-09-26 (gate: production's KV cache back to f16 — it had drifted to q8_0 since 2026-08-08 #386). ADR 0005's per-model safety net never applied, because no model had an override.OLLAMA_KV_CACHE_TYPE, which gives the f16 default.q8_0changes think-on results. In the v0.34.4 fold's think-on protocol on gfx1151, two qwen3.6 grounding cases never finished at 131072 underq8_0. Captured cold, the KV type moves the trajectory within a few hundred characters, and f16 makes the remaining loops start later: one qwen3.6 case from about token 14,650 to about 23,000. No cold case turns from a loop into a finish on f16 alone. (Corrected: this bullet first said one case finishes only with f16. Cold, it finishes underq8_0too, in 17,626 tokens, and its loop in the protocol came from the run's state.)What changes
ADR 0043 sets the policy:
kv_cache_type);ADR 0005's server-wide
q8_0default is marked superseded.tasks/kv-precision-think-loops.md holds the maintainer's question and the protocol. It compares KV
{q8_0, f16, f32}against flash attention{on, off}, as cold captures of each host's looping cases. It has the gfx1151 results and sections for CUDA and Metal.Tools in
vision-suite/:thinkcap.pymakes one cold capture with its thinking kept, and now records which case it captured;kvloop.shruns the arms on a docker host and logs the runner's--cache-type-k/vand--flash-attnflags;kvloop_read.pyreads out the loop profile and the suite's score.cmp_scored.pyshows which quality fields moved between two score files, test by test, and which side each move favours.ADR 0044 and SPEC H25 drop f32 with flash attention on from every KV test. It is f16, byte for byte: on gfx1151, 143,475 characters of qwen3.6 thinking; on CUDA, all six gemma4:26b pairs.
fattn.cuconverts f32 K/V to f16 before its kernels. f32 with flash attention off stays a precision diagnostic. One pair is re-checked when the llama.cpp pin moves.kvloop.shnow defaults tof16:1 f16:0 f32:0.The
kv-loop-checkskill (.claude/skills/) holds what is settled, the procedure and the traps, for any host's agent.Pointers: the per-model spec's "Recommended deployment", the vision-suite README and the fork table in
README.mdnow point at ADR 0043.promptcap.py(prompt variants and card sampling) joins the suite.One mechanism to know before reading f32 results. At b11081, CUDA's and HIP's flash attention convert an f32 K/V cache to f16 before their kernels run (
fattn.cu). So f32 with flash attention on should reproduce f16 exactly, and only f32 with flash attention off makes the attention itself more precise.Production's KV cache per platform (2026-09-28)
amd-server)--cache-type-k f16 --cache-type-v f16. The compose file carries it too (MaxusAI/ollama-deployments31923a9)ai-server/mlx-cuda)OLLAMA_KV_CACHE_TYPE=f16. All 9 KV allocations in its post-deploy preflight are f16 (#390)mlx-metal)What the other hosts answered
kvloop.shon gemma4:26b, on both builds, 24 captures (task doc, CUDA section).bbox_contract_real_1imgonce, at 65536. That case loops on every gfx1151 path.gfx1151 results
Where the loop starts (estimated token), or how the case ends:
q8_0, FA onbbox_contract_real_1imgbbox_contract_adv_realbbox_contract_real_1imgThe reading:
adv_norm1in 1,395 tokens).f32 with flash attention on is byte-identical to f16 on gfx1151 too (ADR 0044).
The prompt and sampling check found the loop's trigger. It is the prompt's one unanswerable sentence: "give the size YOU used" after an internal resize the model cannot see.
So the fix for these loops is the prompt. SPEC C1 (pin norm-1000) already says so, and qwen3.6 answers in 0–1000 whatever the prompt says. At production's sampling, none of the six runs loops.
The KV type against the fold's own change (the fold's think-on protocol, qwen3.6, greedy, full ladder). The
fold2parm has now finished under both KV types. Of the 25 tests that finish in both runs, the KV type moves the quality of 20: 12 toward f16 and 8 towardq8_0. The fold's structured-output change (foldagainstfold2p, bothq8_0) moves 4. The largest moves go both ways: IoU 0.057 → 0.855 with f16 onbboxm_free_noanc_named, and 0.957 → 0.079 onbboxm_free_noanc_pos.adv_realfinishes only under f16. So f16 is not a quality setting, and two arms or two hosts compare only at one KV type (ADR 0043, new consequence).A second trap,
multi_3img_anchored(#375 open item 8, GGUF leg on gfx1151, gemma4:26b, f16). The anchored prompt loops from about token 4,148 and re-derives image 1's height. With only its "size YOU used" sentence replaced, it finishes in 5,875 tokens (size) or 5,633 (commit), with every question right. The control,multi_3img, finishes in 3,616. Underq8_0the same case finished in 8,331 tokens, so here f16 is the side that loops: the numerical path moves where a loop starts, in both directions, and the prompt sets the trap.Verification
mainat15e8e08eb. ADR numbers 0043 and 0044 and SPEC H25 are free there.check_source_paths.py --changed-since origin/mainis clean.test_summarizers.py(includingcmp_scored.py's tests) andtest_rescore.pypass.kvloop_read.pyreads every gfx1151 capture, including the older ones that have no capture block, and the multi-image captures.bash -n kvloop.shis clean.amd-server/rocm-gfx1151🤖 Generated with Claude Code