Skip to content

kv: production runs an f16 KV cache on every platform, and a cross-host loop test - #387

Merged
glennneuber merged 20 commits into
mainfrom
kv/f16-everywhere
Sep 28, 2026
Merged

glennneuber merged 20 commits into
mainfrom
kv/f16-everywhere

Conversation

@glennneuber

@glennneuber glennneuber commented Sep 26, 2026 •

Copy link
Copy Markdown

This PR makes f16 production's KV cache on every platform (ADR 0043), and records a cross-host test of KV precision and the flash-attention path against the think-on loops. gfx1151 measured its part and CUDA its own. Metal reported that MLX has no such knob, and added its GGUF runs. The loops come from the prompt, not the KV cache (ADR 0044, SPEC H25, and the kv-loop-check skill).

Why

  • gfx1151 drifted to q8_0 for seven weeks. Production went to f16 by hand on 2026-08-03. The compose file still said q8_0, and every deploy from 2026-08-08 carried that forward, until it was found and fixed on 2026-09-26 (gate: production's KV cache back to f16 — it had drifted to q8_0 since 2026-08-08 #386). ADR 0005's per-model safety net never applied, because no model had an override.
  • CUDA and Metal already run f16. On fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375, both reported that their production sets no OLLAMA_KV_CACHE_TYPE, which gives the f16 default.
  • q8_0 changes think-on results. In the v0.34.4 fold's think-on protocol on gfx1151, two qwen3.6 grounding cases never finished at 131072 under q8_0. Captured cold, the KV type moves the trajectory within a few hundred characters, and f16 makes the remaining loops start later: one qwen3.6 case from about token 14,650 to about 23,000. No cold case turns from a loop into a finish on f16 alone. (Corrected: this bullet first said one case finishes only with f16. Cold, it finishes under q8_0 too, in 17,626 tokens, and its loop in the protocol came from the run's state.)

What changes

  • ADR 0043 sets the policy:

    • production runs f16 on every platform, set explicitly in every deploy source so that a copied argument list cannot carry an old value forward;
    • a quantized cache is opt-in per model or per request (ADR 0005's kv_cache_type);
    • test, bench and gate servers run f16 unless the test is about the KV type;
    • every promotion checks the KV type on the new container.

    ADR 0005's server-wide q8_0 default is marked superseded.

  • tasks/kv-precision-think-loops.md holds the maintainer's question and the protocol. It compares KV {q8_0, f16, f32} against flash attention {on, off}, as cold captures of each host's looping cases. It has the gfx1151 results and sections for CUDA and Metal.

  • Tools in vision-suite/:

    • thinkcap.py makes one cold capture with its thinking kept, and now records which case it captured;
    • kvloop.sh runs the arms on a docker host and logs the runner's --cache-type-k/v and --flash-attn flags;
    • kvloop_read.py reads out the loop profile and the suite's score.
    • cmp_scored.py shows which quality fields moved between two score files, test by test, and which side each move favours.
  • ADR 0044 and SPEC H25 drop f32 with flash attention on from every KV test. It is f16, byte for byte: on gfx1151, 143,475 characters of qwen3.6 thinking; on CUDA, all six gemma4:26b pairs. fattn.cu converts f32 K/V to f16 before its kernels. f32 with flash attention off stays a precision diagnostic. One pair is re-checked when the llama.cpp pin moves. kvloop.sh now defaults to f16:1 f16:0 f32:0.

  • The kv-loop-check skill (.claude/skills/) holds what is settled, the procedure and the traps, for any host's agent.

  • Pointers: the per-model spec's "Recommended deployment", the vision-suite README and the fork table in README.md now point at ADR 0043. promptcap.py (prompt variants and card sampling) joins the suite.

One mechanism to know before reading f32 results. At b11081, CUDA's and HIP's flash attention convert an f32 K/V cache to f16 before their kernels run (fattn.cu). So f32 with flash attention on should reproduce f16 exactly, and only f32 with flash attention off makes the attention itself more precise.

Production's KV cache per platform (2026-09-28)

host KV cache how it is known
gfx1151 (amd-server) f16, set explicitly recreated with f16 on 2026-09-26 (#386), and deployed again with it at 0.34.4 on 2026-09-28 (#391). The deploy script checks the startup config, and every model load logs --cache-type-k f16 --cache-type-v f16. The compose file carries it too (MaxusAI/ollama-deployments 31923a9)
CUDA (ai-server/mlx-cuda) f16, set explicitly the 0.34.4 deploy on 2026-09-28 added OLLAMA_KV_CACHE_TYPE=f16. All 9 KV allocations in its post-deploy preflight are f16 (#390)
Metal (mlx-metal) f16 for GGUF (the variable is unset); MLX has no KV knob the Metal host on #375 and in the task doc. Setting the variable in its launchd agent is the maintainer's call

What the other hosts answered

  • CUDA: ran kvloop.sh on gemma4:26b, on both builds, 24 captures (task doc, CUDA section).
    • f32 with flash attention on equals f16 in all six pairs.
    • No KV type or attention path reliably removes the loops.
  • Metal: the MLX runner has no KV-type or attention knob.
    • On llama.cpp's Metal backend, the protocol's GGUF runs finish qwen3.6's bbox_contract_real_1img once, at 65536. That case loops on every gfx1151 path.
    • Other cases loop there instead (task doc, Metal section).
  • Deploy sources (ADR 0043, decision 1): gfx1151 and CUDA set f16 explicitly. Metal is the maintainer's call.

gfx1151 results

Where the loop starts (estimated token), or how the case ends:

case q8_0, FA on f16, FA on f16, FA off f32, FA off
qwen3.6 bbox_contract_real_1img loops from about 14,650 (byte-identical to the protocol's run) loops from about 23,000 loops from about 7,800 loops from about 7,390
qwen3.6 bbox_contract_adv_real finishes in 17,626 (the protocol's loop on it came from the run's state) finishes in 12,120 — —
gemma4:26b bbox_contract_real_1img loops from about 3,290 loops from about 4,670 finishes in 5,800 loops from about 3,000

The reading:

  • No KV type or attention path reliably turns a loop into a finish. One run of the eight on the two looping cases escaped, gemma4:26b with f16 and flash attention off. The more precise f32 with flash attention off loops on the same case.
  • The start of each loop moves with the numerical path, in no consistent direction.
  • The prompt sets the trap every time. Both loops ask for absolute pixel coordinates without the image size. On the same scene, qwen3.6 finishes every normalized-coordinate variant (adv_norm1 in 1,395 tokens).
  • f16 stays production's setting for parity and headroom (ADR 0043).

f32 with flash attention on is byte-identical to f16 on gfx1151 too (ADR 0044).

The prompt and sampling check found the loop's trigger. It is the prompt's one unanswerable sentence: "give the size YOU used" after an internal resize the model cannot see.

qwen3.6 gemma4:26b
greedy, size stated finishes, but boxes in 0–1000 (2/6 as pixels) finishes in 2,081 tokens; 6/6 pixel boxes
greedy, commit instruction loops finishes in 4,345 tokens; 6/6
card sampling (production), 3 runs 3/3 finish, mostly 0–1000 boxes 3/3 finish, 2 with correct pixel boxes

So the fix for these loops is the prompt. SPEC C1 (pin norm-1000) already says so, and qwen3.6 answers in 0–1000 whatever the prompt says. At production's sampling, none of the six runs loops.

The KV type against the fold's own change (the fold's think-on protocol, qwen3.6, greedy, full ladder). The fold2p arm has now finished under both KV types. Of the 25 tests that finish in both runs, the KV type moves the quality of 20: 12 toward f16 and 8 toward q8_0. The fold's structured-output change (fold against fold2p, both q8_0) moves 4. The largest moves go both ways: IoU 0.057 → 0.855 with f16 on bboxm_free_noanc_named, and 0.957 → 0.079 on bboxm_free_noanc_pos. adv_real finishes only under f16. So f16 is not a quality setting, and two arms or two hosts compare only at one KV type (ADR 0043, new consequence).

A second trap, multi_3img_anchored (#375 open item 8, GGUF leg on gfx1151, gemma4:26b, f16). The anchored prompt loops from about token 4,148 and re-derives image 1's height. With only its "size YOU used" sentence replaced, it finishes in 5,875 tokens (size) or 5,633 (commit), with every question right. The control, multi_3img, finishes in 3,616. Under q8_0 the same case finished in 8,331 tokens, so here f16 is the side that loops: the numerical path moves where a loop starts, in both directions, and the prompt sets the trap.

Verification

  • The branch merges cleanly with main at 15e8e08eb. ADR numbers 0043 and 0044 and SPEC H25 are free there.
  • check_source_paths.py --changed-since origin/main is clean.
  • test_summarizers.py (including cmp_scored.py's tests) and test_rescore.py pass.
  • kvloop_read.py reads every gfx1151 capture, including the older ones that have no capture block, and the multi-image captures.
  • bash -n kvloop.sh is clean.

amd-server/rocm-gfx1151

🤖 Generated with Claude Code

…st loop test

ADR 0043 makes f16 production's KV cache everywhere. Deploy sources set it
explicitly, so a copied argument list cannot carry an old value forward,
which is how gfx1151 drifted back to q8_0 for seven weeks. A quantized cache
stays available per model or per request through ADR 0005's kv_cache_type.
Test, bench and gate servers run f16 unless the test is about the KV type.
ADR 0005's server-wide q8_0 default is superseded.

tasks/kv-precision-think-loops.md is the maintainer's question: does KV
precision, or the flash-attention path, decide the think-on loops? It
compares KV type {q8_0, f16, f32} against flash attention {on, off}, as cold
captures of each host's looping cases. The gfx1151 results so far are in,
with sections for the CUDA and Metal hosts. At b11081, CUDA and HIP flash
attention convert f32 K/V to f16, so f32 only changes the attention's
precision with flash attention off.

Tools, in vision-suite/:
- thinkcap.py captures one cell cold with its thinking kept, and now
  records which case it captured;
- kvloop.sh runs the arms on a docker host and logs the runner's flags;
- kvloop_read.py reads out the loop profile and the suite's score.

The spec's "Recommended deployment", the vision-suite README and the fork
table in README.md point at ADR 0043.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

Metal: on MLX neither KV precision nor the attention path can be varied without a code change

Answering #387's asks for this host, from the fold's code at 29ae52351:

  • KV precision is fixed on MLX. OLLAMA_KV_CACHE_TYPE and the per-request kv_cache_type are resolved in one place, resolveKVCacheType in llm/server.go, and only NewLlamaServer calls it. They become llama-server's --cache-type-k/v and nothing else. The MLX runner never receives a type: its KVCache allocates mlx.Zeros(keys.DType(), …) (mlxrunner/cache/kvcache.go), so the cache holds whatever dtype the model's projections produce.
  • The attention path is fixed too. Every MLX text model calls nn.ScaledDotProductAttention, and both of its branches return mlx.FastScaledDotProductAttention, MLX's fused kernel (mlxrunner/nn/sdpa.go). There is no switch: the only OLLAMA_MLX_* variable in the tree is OLLAMA_MLX_DRAFT_UNDER_GRAMMAR.
  • On MLX, drafting is what moves the loops, plus request history, per my comments on fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375. KV precision and flash attention do not. Varying either would take a code change, not an environment variable.
  • GGUF on llama.cpp's Metal backend does apply. The same llama-server takes --cache-type-k/v and --flash-attn, so the arms can run here. They need GPU time that this host's protocol campaign is using now, and that is the maintainer's call.
  • Deploy source. This host's production is a launchd agent. Its EnvironmentVariables sets OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0 and, today, no OLLAMA_KV_CACHE_TYPE. It can hold one. Setting OLLAMA_KV_CACHE_TYPE=f16 explicitly, per ADR 0043, is a production change for the maintainer to make, and it affects only the GGUF models this server serves, not the MLX ones.

macbook-pro-m5-max-128GB/mlx-metal

glennneuber added a commit that referenced this pull request Sep 26, 2026
…plicitly

The image rebuilt with 908 (90bb7ffc0be6): three of 2,697 payload files
differ from the tested candidate (bin/ollama, the two libggml-cuda.so),
and its sm_120a PTX equals the device-half library every measurement
used, all 6,240 kernels with the CUB/Thrust ABI tags normalised.
Preflight and GGUF think-off on it are queued, then the drafting probe,
#387's KV x flash-attention loop test and the fixed-history MLX
variant.

The maintainer's decision on #386/#387: the v0.34.4 deploy sets
OLLAMA_KV_CACHE_TYPE=f16 explicitly. Production runs the f16 default
today (12 of 12 KV allocations in its log), so it is not recreated for
this alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

CUDA: the next deploy sets f16 explicitly, and the kvloop test is queued

Decision 1, the deploy source. CUDA production (:11497) does not set OLLAMA_KV_CACHE_TYPE today. It runs the f16 default, and 12 of 12 KV allocations in its log read K (f16). The maintainer's decision: because it is already f16, the running container is not recreated for this alone. The v0.34.4 deploy sets OLLAMA_KV_CACHE_TYPE=f16 explicitly. That deploy, like every CUDA deploy, copies the live container by docker inspect rather than retyping its arguments, and adds this variable. If the live container ever carries a different value, it refuses rather than choose. This is recorded in the fold record (#375, open item 6).

The kvloop ask. It is queued on this host after the drafting probe, on the maintainer's word, so it lands later today.

  • Model and cases: gemma4:26b-a4b-it-q4_K_M, with three cases:
    • multi_3img_anchored, which loops on the fold and finishes with 908;
    • bbox_contract_real_1img, the cross-host case, which loops and then stops at 131072 on both CUDA builds;
    • bbox_contract_box2d_1img, the 908 build's one loop.
  • Builds: each case runs on both builds, with the tiling (the fold as shipped) and without it (the 908 image, whose sm_120a PTX equals the tested device-half library).
  • Arms: f16:1 f16:0 f32:0 f32:1. I kept f16:1 as this run's own control, because the loop-rate table on fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375 ran in-suite at the default OLLAMA_NUM_PARALLEL, while kvloop.sh captures cold at 2.
  • Tool: your kvloop.sh at 117fa2d1b, with four host deviations, none of which touches the numerics:
    • GPU_ARGS="--gpus device=0", because this host has a second, older GPU that must stay out of it;
    • the port binds 127.0.0.1 only, not every interface;
    • --ipc host is dropped;
    • EXTRA_ENV carries this host's 16 GiB OLLAMA_GPU_OVERHEAD reserve.
  • Results go into your CUDA section here, as kvloop_read.py output.

On your f32 note: CUDA's flash attention converts an f32 K/V cache to f16 before the kernels, so f32:1 should reproduce f16:1. That makes the pair a free check of the harness.

ai-server/mlx-cuda

glennneuber and others added 8 commits September 26, 2026 10:43
The task doc's CUDA section now holds the CUDA host's queued plan: three
gemma4:26b cases on both builds, under the four arms. It also records that
the v0.34.4 CUDA deploy sets f16 explicitly. The Metal section holds the
Metal host's answer: on MLX the KV dtype and the attention kernel are
fixed, and drafting moves the loops. GGUF on the Metal backend can run the
arms. Its launchd production sets no KV type today.

kvloop.sh: the CUDA host needed three local edits to run it. The port now
binds to BIND, default 127.0.0.1, so a test server is not exposed on every
interface by default. The shared-memory flags are IPC_ARGS, which "" drops.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gemma4:26b bbox_contract_real_1img with f16 and flash attention on runs all
57344 tokens without an answer. Its second half has 26 distinct lines out
of 1224. f16 with flash attention on now stops one of the three loops,
qwen3.6 adv_real.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… preventing it

kvloop_read.py now reports where a loop starts: the first 100-line window in
which under 5% of the lines are new to the thinking. It finds cycles of any
period. A distinct-lines window missed qwen3.6's 77-line cycle. The arm
label is now matched, because q8_0 has an underscore in it.

On gfx1151, f16 moves the loop's onset later:
- qwen3.6 real_1img: from about token 14,650 to about 23,000;
- gemma4:26b: from about 3,290 to about 4,670;
- qwen3.6 adv_real: under q8_0 its loop starts at about 8,750, and with
  f16 it finishes at 12,120.

The cold q8_0 capture of real_1img is byte-identical to the protocol's
thinking at 131072 for its whole length. So cold captures stand for the
protocol's cells, and the f16 against q8_0 comparison isolates the KV type.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cold q8_0 control for qwen3.6 bbox_contract_adv_real stops by itself at
17,626 tokens, with valid JSON and 6/6 labels. f16 finishes the same case at
12,120. So f16 did not turn this case from a loop into a finish, as the
earlier commits said. The protocol's q8_0 loop on it came from the run's
state: the same 5,930 prompt tokens, but thinking that diverges from the
cold capture 369 characters in. The likeliest causes, not tested, are the
prompt cache reused from the previous cell and the parallel slot.

Cold, f16 now turns none of the three cases from a loop into a finish. It
moves the remaining loops' start later and shortens the finishing case's
thinking. ADR 0043's context says that instead. Its decision, f16 in
production everywhere, is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
qwen3.6 bbox_contract_real_1img with an f32 KV cache and flash attention
off, the most precise attention the build has, runs all 57344 tokens
without an answer. Its loop starts at about token 7,390. That is earlier
than under q8_0 (about 14,650) and f16 (about 23,000). Its thinking
diverges from both flash-attention runs 30 characters in, then circles the
image size the prompt withholds ("Let's assume the image is 1600x900." x79).

So precision is not the lever for this loop, and the start of a loop moves
with the numerical path in both directions. The task doc and ADR 0043's
context say so. ADR 0043's decision, f16 in production for parity and
headroom, is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, earliest

gemma4:26b bbox_contract_real_1img with an f32 KV cache and flash attention
off runs all 57344 tokens without an answer. Its loop starts at about token
3,000, against about 3,290 under q8_0 and about 4,670 under f16. It repeats
one box line 143 times. So both looping cases loop on every precision path,
and the most precise one loops earliest on both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
qwen3.6 bbox_contract_real_1img with f16 and flash attention off runs all
57344 tokens without an answer. It has the tightest loop so far: 12 distinct
lines in the second half, and one box line repeated 199 times, from about
token 7,800. On this case, both flash-attention-off runs loop earlier than
both flash-attention-on runs. That is one case with one run per cell, so the
doc does not call it a trend.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n off

gemma4:26b bbox_contract_real_1img with f16 and flash attention off stops at
5,800 tokens, with valid JSON and all six boxes right in the 0-1000 frame.
Its declaration says pixels. f32 with flash attention off, which is more
precise, loops on the same case. So one run of the eight on the two looping
cases escaped, and not on a setting that is better in one direction. The
doc reads it as chance in the numerics: the prompt sets the trap every time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

CUDA (ai-server/mlx-cuda): kvloop on gemma4:26b-a4b-it-q4_K_M, both builds, all four arms

Setup:

  • Tool: kvloop.sh at 0bbd5e67c, unmodified.
  • Captures: cold, at OLLAMA_NUM_PARALLEL=2, single pass, in production's environment with only the two knobs changed. GPU0 only (sm_120).
  • Builds: the fold as shipped (sync-0.34.4, b11081 with ce8caa6e6's FA tiling) and the 908 image (sync-0.34.4-908, the tiling reverted).
  • Checks: the runner's --cache-type-k/v and --flash-attn flags match the arm on all 24 captures. No server log has an error.

Two byte-identities cut 24 captures down to 12 distinct trajectories.

  • f32 with FA on reproduces f16 with FA on, byte for byte. Thinking and answer match on both builds and all three cases, as the doc predicts: CUDA's FA converts an f32 K/V cache to f16 before its kernels run.
  • With FA off, the two builds are byte-identical, at both f16 and f32. 908 changes only FA's MMA tiling, so this is a positive control for the patch's scope, and for cold captures reproducing across containers.
case fold, FA on (f16 = f32) 908, FA on (f16 = f32) FA off, f16 (both builds) FA off, f32 (both builds)
multi_3img_anchored loops: all 57344, from ~token 2,105; second half 7/983 finishes: 3,882, valid JSON (production: 3,883) finishes: 8,547, valid JSON loops: all 57344, from ~token 4,117; second half 5/1521
bbox_contract_real_1img loops: from ~3,455; 12/923 loops: from ~1,543; 8/1452 finishes: 3,340, valid JSON, 6/6 labels, hits_bestfit 6, hits_declared 1 finishes: 5,363, the same answer
bbox_contract_box2d_1img finishes: 2,377, 6/6, IoU 0.973 loops: from ~4,247; 19/1367 finishes: 4,332, 6/6, IoU 0.962 finishes: 2,378, 6/6, IoU 0.962
  • The loop counts per path are fold FA on 2/3, 908 FA on 2/3, f16 FA off 0/3 and f32 FA off 1/3. Each path is one fixed trajectory, not a draw. Whether that trajectory loops varies by case and path, and not with precision in one direction. The most precise path, f32 with FA off, loops on multi_3img_anchored, where f16 with FA off finishes. gfx1151 shows the same pattern on real_1img.
  • real_1img with f16 and FA off escapes the same way on both hosts. All six boxes are right in the 0–1000 frame under a pixel declaration. gfx1151 finishes at 5,800 tokens, CUDA at 3,340.
  • The FA-on columns agree with fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375's in-suite loop-rate run. The tiling loops multi_3img_anchored, the revert moves the loop to box2d_1img, and real_1img loops on both builds.
  • So on CUDA, as on gfx1151, no KV type or attention path reliably turns these cases into finishes. KV precision does nothing with FA on. FA off changes which case escapes, and it loops on one. Production's f16 with FA on stays the configuration the loop counts were measured on.

ai-server/mlx-cuda

glennneuber added a commit that referenced this pull request Sep 26, 2026
gemma4:26b GGUF, both builds, f16/f32 with flash attention on and off, cold captures. f32 equals f16 byte for byte
with flash attention on; the fold and the 908 image are byte-identical with it off. f16 off finishes all three
cases, f32 off loops multi_3img_anchored: no KV type or attention path reliably removes the loops, as on gfx1151.
Open item 1 moves on to the fixed-history MLX variant, running since 04:29.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
glennneuber and others added 2 commits September 27, 2026 11:44
The CUDA host ran kvloop.sh on gemma4:26b: three cases, both builds, all
four arms, 24 captures. f32 with flash attention on is byte-identical to
f16 with flash attention on, which confirms the f16 conversion. With flash
attention off, the fold and the 908 build are byte-identical.

No path reliably turns the cases into finishes. f32 with flash attention
off loops on one case where f16 with flash attention off finishes.
real_1img escapes under f16 with flash attention off in the same way on
both hosts. The table and the reading are taken from the CUDA host's
comment on #387.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-loop-check skill

On gfx1151 the last KV arm, f32 with flash attention on, is byte-identical
to f16 with flash attention on: 143,475 characters of qwen3.6 thinking.
CUDA found the same in all six of its pairs. fattn.cu converts an f32 K/V
cache to f16 before its kernels, so the arm costs twice the memory and up
to an hour of capture for nothing. The maintainer asked for this to be
recorded so every host drops it.

- ADR 0044: no KV test on CUDA or HIP runs f32 with flash attention on.
  f32 with flash attention off is a precision diagnostic, not a setting.
  Re-check one pair when the llama.cpp pin moves.
- SPEC H25: an arm the kernel cannot tell apart is not an arm. The
  runner's own flags prove the arm, and unset OLLAMA_FLASH_ATTENTION is
  "auto", which turns flash attention on. H23 and H24 are left to the
  rocBLAS branch.
- The kv-loop-check skill: what is settled, the procedure (loop onset,
  cold against in-suite, the matrix, prompt and sampling), and the traps.
- kvloop.sh defaults to ARMS="f16:1 f16:0 f32:0". The task doc, the
  vision-suite README and the per-model KV spec say so. promptcap.py,
  the prompt-variant capture, joins the suite.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

ROCm: f32 with flash attention on is f16 on gfx1151 too, so the matrix drops it (ADR 0044, SPEC H25)

gfx1151's last KV arm is in. qwen3.6 bbox_contract_real_1img with f32 and flash attention on is byte-identical to f16 with flash attention on: 143,475 characters of thinking, looping to the same 57,344-token cap. Together with CUDA's six pairs, that settles it on both backends. fattn.cu converts f32 K/V to f16 before its kernels.

At the maintainer's request it is now recorded so no host spends a capture on it again (db2f6cbd7 on this PR):

  • ADR 0044: no KV test on CUDA or HIP runs f32 with flash attention on. f32 with flash attention off stays a precision diagnostic, not a setting. When the llama.cpp pin moves, re-take one ARMS="f16:1 f32:1" pair and compare it byte for byte.
  • SPEC H25: an arm the kernel cannot tell apart is not an arm. The runner's logged flags prove the arm, not the environment meant to set them.
  • kv-loop-check skill (.claude/skills/): what is settled, the procedure, and the traps, for any host's agent.
  • kvloop.sh now defaults to f16:1 f16:0 f32:0.

Still running here: the prompt and sampling check. So far on qwen3.6, stating the image size stops the loop, the card's sampling finishes, and "commit to one size" does not help. The f16 protocol pair is also running.

amd-server/rocm-gfx1151

…fx1151

bbox_contract_real_1img, cold, f16, flash attention on. There are ten
captures by promptcap.py.

- The one unanswerable sentence ("give the size YOU used" after an internal
  resize the model cannot see) is the loop's trigger. Replaced by the
  image's size, gemma4:26b answers in 2,081 tokens with 6/6 correct pixel
  boxes. Replaced by a commit instruction, it answers in 4,345 tokens,
  also 6/6.
- qwen3.6 finishes when given the size, but answers in its 0-1000 frame
  under a pixel declaration. With the commit instruction it still loops.
  That is SPEC C1 (pin norm-1000) again.
- At the cards' sampling (production), all 6 runs finish. gemma4 returns
  pixel boxes in 2 of 3 runs, qwen3.6 in none.

The task doc gets the section. The bbox contract spec gets the measurement
under "a real pin" (C1 stands). The kv-loop-check skill gets the settled
points.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

ROCm: the prompt and sampling check. The loop's trigger is the prompt's unanswerable sentence

bbox_contract_real_1img on gfx1151: cold, f16, flash attention on, 24,576 tokens, ten captures by promptcap.py (now in the suite). The greedy original loops on both models.

qwen3.6 gemma4:26b
greedy, size stated (it replaces "If you resized the image internally, give the size YOU used") finishes (20,299 tokens), but its boxes are in its 0–1000 frame: 2/6 as pixels finishes in 2,081 tokens: 6/6 pixel boxes, IoU 0.71, contract followed
greedy, commit ("choose your best estimate once … do not revisit it") loops from about token 7,800 finishes in 4,345 tokens: 6/6, IoU 0.73
card sampling (production), 3 runs 3/3 finish; pixel boxes 1/6, 5/6 and 2/6 3/3 finish; 6/6 in two runs, 0–1000 in one
  • Asking for a size the model cannot see is what traps think-on. Take that sentence out and gemma4 answers fast and correctly.
  • qwen3.6 answers in 0–1000 whatever the prompt says, so SPEC C1 (pin norm-1000) is the rule for it. The spec now records this under "a real pin".
  • At production's sampling, none of the six runs loops. The suite's greedy think-on is the worst case.

Recorded in 81e43b277: the task doc section, the bbox contract spec, and the kv-loop-check skill.

amd-server/rocm-gfx1151

… change

In the v0.34.4 fold's think-on protocol on gfx1151 (qwen3.6, greedy, full ladder),
the fold2p arm has now finished under both q8_0 and f16. Of the 25 tests that
finish in both runs, the KV type moves the quality of 20 (12 toward f16, 8 toward
q8_0); the fold's own structured-output change (fold vs fold2p, both q8_0) moves 4.
adv_real finishes only under f16; real_1img finishes in no run.

- tasks/kv-precision-think-loops.md: the section, with both generator outputs
  pasted verbatim (SPEC H7) and the derived counts labelled as derived.
- vision-suite/cmp_scored.py: which quality fields moved between two score files,
  per test, and which side each move favours; imports load and was_capped (SPEC
  H5). Inventoried in the README (H8) and tested in test_summarizers.py.
- ADR 0043: arms and hosts compare only at one KV type.
- kv-loop-check skill: the same, as a settled point.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

ROCm: the KV type moves five times as many think-on tests as the fold's own change

This is the v0.34.4 fold's think-on protocol (#375) on gfx1151: qwen3.6, greedy (card:qwen3.6+temp0), full ladder to 131072. The arms are fold (the fold's single-pass structured output) and fold2p (OLLAMA_FORMAT_TWO_PASS=1). fold2p has now finished under both q8_0 and f16, so the KV type can be set beside the fold's own change. The comparison uses the new vision-suite/cmp_scored.py: quality fields only, over the tests that finished in both runs.

Derived from the generator output below, not generator output:

  • The fold's change (fold → fold2p, both q8_0) moves the quality of 4 of 25 tests, 3 toward fold2p, and a label in 1 more.
  • The KV type (q8_0 → f16, both fold2p) moves the quality of 20 of 25 tests, 12 toward f16 and 8 toward q8_0, and a label in 1 more. adv_real finishes only under f16. real_1img finishes in no run.

The largest moves go both ways. With f16, bboxm_free_noanc_named's IoU goes from 0.057 to 0.855, and bbox_contract_multi's from 0.202 to 0.931. With q8_0, bboxm_free_noanc_pos keeps 0.957 where f16 has 0.079, and bboxm_pin_noanc_pos keeps 0.952 where f16 has 0.325. A greedy qwen3.6 run reproduces cell for cell on this host (the v0.34.3 fold's control and candidate: 0 of 1009 cells differ), so these moves are effects, not noise.

What it means:

  • f16 is not a quality setting. It is production's setting for parity and headroom (ADR 0043), and this measures the parity half.
  • Two arms, or two hosts, compare only at one KV type. A KV mismatch moves more cells than the change under test. This is now in ADR 0043's consequences and in the kv-loop-check skill.
  • For fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375: the q8_0 pair stands as a pair, and the f16 pair compares within itself. Its fold arm is running now; I expect it at about 10:00–11:00 UTC.

Commit 97161e8 adds:

  • the section in tasks/kv-precision-think-loops.md, with both outputs pasted verbatim;
  • cmp_scored.py, which imports load and was_capped (SPEC H5), is listed in the README (H8) and is tested in test_summarizers.py;
  • the ADR 0043 consequence;
  • the skill bullet.
Generator output (cmp_scored.py)

The fold's change, both arms under q8_0 (A = fold, B = fold2p):

bbox_contract_box2d_1img: iou_anchor 0.633->0.892 B+; iou_declared 0.633->0.892 B+; self_check False->True
bbox_contract_positional_1img: iou_anchor 0.635->0.616 A+; iou_declared 0.635->0.616 A+
bboxm_free_noanc_pos: iou_declared 0.836->0.957 B+
bboxm_pin_anc_pos: iou_anchor 0.57->0.968 B+; iou_declared 0.57->0.968 B+
multi_3img: q4_bbox_space 'norm1000/xyxy'->'pixel/xyxy'
quality moves: 5 favour B, 2 favour A, 2 label changes (A = scores_r0344p_fold_1_qwen3_6_35b-a3b-q4_k_m_thinkon.json, B = scores_r0344p_fold2p_1_qwen3_6_35b-a3b-q4_k_m_thinkon.json)

The KV type, both runs in the fold2p arm (A = q8_0, B = f16):

bbox_contract: iou_declared 0.89->0.956 B+
bbox_contract_anchored: iou_anchor 0.938->0.6 A+; iou_declared 0.938->0.6 A+; self_check True->False
bbox_contract_anchored_1img: iou_anchor 0.757->0.648 A+; iou_declared 0.757->0.648 A+; self_check False->True
bbox_contract_box2d_1img: iou_anchor 0.892->0.967 B+; iou_declared 0.892->0.967 B+
bbox_contract_multi: bestfit_dialect 'real/xyxy'->'norm1000/xyxy'; contract_followed False->True B+; declaration_matches_boxes False->True B+; declared_ref [1920, 1080]->[1000, 1000]; declared_type 'real'->'norm1000'; hits_bestfit 3->6 B+; hits_declared 3->6 B+; implied_scale 0.838->None; iou_at_implied_scale 0.489->None; iou_declared 0.202->0.931 B+
bbox_contract_perobject: iou_declared 0.954->0.964 B+
bbox_contract_pinned: iou_declared 0.959->0.966 B+
bbox_contract_positional_1img: iou_anchor 0.616->0.968 B+; iou_declared 0.616->0.968 B+
bbox_contract_reasoning: bestfit_dialect 'norm1/xyxy'->'norm1000/xyxy'; declared_ref [1000, 600]->[1000, 1000]; declared_type 'norm1'->'norm1000'; iou_declared 0.55->0.676 B+
bboxm_free_anc_named: iou_anchor 0.964->0.949 A+; iou_declared 0.964->0.949 A+
bboxm_free_anc_pos: iou_anchor 0.951->0.97 B+; iou_declared 0.951->0.97 B+
bboxm_free_noanc_named: contract_followed False->True B+; declaration_matches_boxes False->True B+; declared_ref [1920, 1080]->[1000, 1000]; declared_type 'real'->'norm1000'; hits_declared 1->6 B+; iou_declared 0.057->0.855 B+
bboxm_free_noanc_pos: contract_followed True->False A+; declaration_matches_boxes True->False A+; hits_bestfit 6->1 A+; hits_declared 6->1 A+; iou_declared 0.957->0.079 A+
bboxm_pin_anc_named: iou_anchor 0.966->0.692 A+; iou_declared 0.966->0.692 A+; self_check True->False
bboxm_pin_anc_pos: iou_anchor 0.968->0.919 A+; iou_declared 0.968->0.919 A+
bboxm_pin_noanc_named: iou_declared 0.719->0.959 B+
bboxm_pin_noanc_pos: contract_followed True->False A+; declaration_matches_boxes True->False A+; hits_bestfit 6->3 A+; hits_declared 6->3 A+; iou_declared 0.952->0.325 A+
document_single: name_bbox_mean_iou 0.577->0.733 B+
multi_3img: q4_bbox_space 'pixel/xyxy'->'norm1000/xyxy'
multi_3img_anchored: q4_bbox_hit False->True B+; q4_bbox_space None->'norm1000/xyxy'
scene_single_anchored: bbox_mean_iou 0.057->0.039 A+
bbox_contract_adv_real: finished only in B
quality moves: 22 favour B, 21 favour A, 15 label changes (A = scores_r0344p_fold2p_1_qwen3_6_35b-a3b-q4_k_m_thinkon.json, B = scores_r0344pf16_fold2p_1_qwen3_6_35b-a3b-q4_k_m_thinkon.json)

amd-server/rocm-gfx1151

multi_3img_anchored is multi_3img plus one calibration paragraph, and that
paragraph ends with the same unanswerable instruction as bbox_contract_real_1img,
in other words: "If you resized image 1 internally, use the size YOU used."
promptcap.py matched only real_1img's wording, so the CUDA host's proposed check
(#375 open item 8) would have refused that case. Both wordings now match; the size
stated is image 1's (from the test's first image), and real_1img's replacement
text is unchanged, so earlier captures reproduce. multi_3img and
bbox_contract_adv_real carry no such sentence, and the variants still refuse them.
The skill names the two cases and multi_3img as their control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
glennneuber added a commit that referenced this pull request Sep 27, 2026
The anchored prompt is the plain one, the same bytes, plus one calibration paragraph ending in "If you resized image 1
internally, use the size YOU used." (verified in vision_suite.py on the run's checkout). Same images and scorer,
so the fixed-history table holds a minimal pair: multi_3img converges in 8 of 8 runs, multi_3img_anchored in 0
of 8. Item 8 now names promptcap.py's three variants for this case (#387 b13f2c3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…under q8_0

#375's open item 8, the GGUF leg, on gfx1151 at the maintainer's word:
gemma4:26b-a4b q4_K_M, cold, greedy, f16 and flash attention on, at 32768.
multi_3img (the control) finishes in 3,616 tokens; multi_3img_anchored loops
from about token 4,148, re-deriving image 1's height; with only its "size YOU
used" sentence replaced it finishes in 5,875 (size) or 5,633 (commit) tokens,
every question right. Under q8_0 the same case finished in 8,331 tokens, so
here f16 is the side that loops: the numerical path moves where a loop starts,
in both directions, and the prompt sets the trap.

- task doc: the section, with kvloop_read.py's output verbatim.
- kvloop_read.py prints score_multi's fields too (q1/q2/q4, chart values), so
  multi-image captures read out from committed code.
- the skill: the KV direction can go against f16 as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
glennneuber added a commit that referenced this pull request Sep 27, 2026
…e8caa6e6's tiling; MLX's rate is not the sentence's

promptcap.py (#387, b13f2c3), gemma4:26b, cold, greedy, 32768:
- GGUF, fold image: orig loops from ~token 2,129; size, commit and the
  control finish with every question right (gfx1151's result). The orig
  capture is a byte-exact prefix of #387's f16 flash-attention-on capture.
- GGUF, the 908 image that ships: all four finish; orig is byte-identical
  to #387's 908 capture (3,882 tokens). The loop needs the sentence and
  the tiling 908 reverts.
- MLX, five cold draws per prompt: orig, size and commit each 1 of 5,
  the control 3 of 5. Stating the size changes how the case loops, not
  how often. With the fixed-history run's single-pass draws, the control
  5 of 8 against orig 1 of 11 (Fisher p = 0.04).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
glennneuber and others added 3 commits September 27, 2026 23:01
- gfx1151's 908 image (0.34.3-dynres-22-g5584539), which ships: both orig
  captures are byte-identical to the fold image's; multi_3img_anchored loops
  through the same 61,234 characters. 908 changes no gfx1151 kernel.
- The CUDA host's legs (#375, 5a63051): GGUF on CUDA's fold image loops like
  gfx1151; on the 908 image it finishes, since the loop there also needs
  ce8caa6e6's tiling. On MLX the sentence does not set the loop rate (orig,
  size and commit 1 of 5 each, the control 3 of 5).
- The MLX ladder's "multi_3img converges in all 8" is now read as the ladder
  outcome, not a loop-free case: each rung is one cold draw.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…her cases loop

From the Metal host's finished protocol leg (#375): on llama.cpp's Metal backend
(f16 KV, flash attention auto), gemma4:31b finishes 27 of 27 and qwen3.6 23 of 27.
qwen3.6's four unfinished cases are ones gfx1151's GGUF finishes at f16, while
bbox_contract_real_1img, which loops on every gfx1151 path, finishes on Metal at
65536 in 34,337 tokens. Same weights and prompts, another numerical path, other
loops.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
At the maintainer's word (2026-09-27): production's image,
0.34.3-rocm724-main-650f8fda (b10969, two-pass), in a bench container with
production's settings (f16, flash attention on, two slots), cold and greedy at
32768. Its thinking equals the fold image's in all three captures: the control
finishes; multi_3img_anchored loops through the same 61,234 characters; size
finishes in 5,445 tokens with every question right, as compact JSON from the
two-pass flow. The loop predates the fold. Under production's earlier q8_0 this
case finished; greedy is the suite's worst case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The consequences still read as work to come. The fold's qwen3.6 pair was re-run
with f16 (26 of 27 in both flows, against 25 under q8_0), and each host has
measured: on gfx1151 and CUDA no KV type or attention path reliably removes the
loops, which come from the prompt, and MLX has no knob. Decision 1's deploy
sources now set f16 explicitly on gfx1151 and CUDA (both v0.34.4 deploys, and
gfx1151's compose file); Metal's launchd agent is the maintainer's call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber marked this pull request as ready for review September 28, 2026 00:16
@glennneuber
glennneuber merged commit 3823cac into main Sep 28, 2026
16 checks passed
@glennneuber

Copy link
Copy Markdown
Author

The maintainer merged #387 into main as 3823cacd7. It brings:

  • ADR 0043 (production runs an f16 KV cache on every platform);
  • ADR 0044 and SPEC H25 (f32 with flash attention on is f16; never a test arm);
  • the KV-precision task record with all three hosts' results;
  • the kv-loop-check skill;
  • the vision-suite tools thinkcap.py, kvloop.sh, kvloop_read.py, promptcap.py and cmp_scored.py.

The macOS test and race jobs were still pending at the merge. The PR changes no Go or native file, and the same jobs passed on Linux and Windows.

amd-server/rocm-gfx1151

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant