docs: ROCm gate — #17459 and #17475 do not reproduce on gfx1151 - #304
Conversation
…fx1151 Answers the brief in docs/maxusai/tasks/rocm-gate-issues-17459-17475.md. ollama#17459: 0 <unused49> tokens in 96 requests across eight payload/environment combinations (b9888; b10091 fork and stock; b10864 fork dio-on, fork dio-off, stock), cold and warm, think both ways. On this host think=false is the arm that works, the opposite of the report. ollama#17475: 0 contaminated victims in 436 extractions (484 including superseded runs) across six configurations and all three protocols, with donor aborts and multi-model swapping verified to have actually occurred. Three instrument corrections were needed first: the 2.5 s abort window (8 s never fired here), continuing victims until the noise loop cycled (V3 had swapped 1-3 times), and a marker that does not collide with "SYNTHETIC TEST DOCUMENT" on the forms. Incidental and more consequential than either issue: 0.32.x reports 27 GiB of iGPU memory where 0.34.x reports 95.4 GiB with the ROCm runtime held constant, so production places 25-26 of gemma4:31b's 61 layers on the GPU where 0.34.1 places 61. Observed, not explained. Direct I/O, which the gate treats as a risk, is also what makes repeated model loading viable here. Neither issue can satisfy its gate clause as written -- both are closed not_planned, so no fix can be "in the target tag". Clause 4 (vision A/B on the candidate) is the remaining piece of work the gate actually asks for. Harness committed under docs/maxusai/tasks/rocm-gate/ so the reproduction can be repeated or disputed; host store path and GPU GIDs are env-supplied. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Update that changes this PR's conclusion: clause 4 has now been run, and the candidate fails it. This PR argues the pin rests on two unreproducible bugs plus a measured cost, and names the vision A/B on Clause 4: FAILFull vision campaign, five models,
Fine-text OCR collapses on the same three (4/4/4 → 0/0/0 on What that does to this documentThe framing needs revising, and I would rather say so here than let the PR merge as written:
The regression, narrowedSubsequent work localized it considerably — details in the comment on #309. Briefly: with the image-token budget pinned per cell, Controls held: the pinned build re-run today reproduces its own baseline exactly (host has not drifted); unaffected by direct I/O; unaffected by layer placement. Mechanism still unknown. Refuted by experiment: fp16 accumulation in the vision tower, multi-sub-batch image decode, and reverting the b10864 resize-algo flip. Suggested dispositionI would rather not merge this as-is. Either it gets revised to carry the clause 4 result, or it merges as a dated record of what was known on 2026-09-17 with a pointer to the regression work — but it should not stand as the current view of whether the gate can lift. 🤖 Generated with Claude Code |
…906 fixes it Clause 4 was run 2026-09-18. Unpatched b10864 dropped qwen3.6:35b-a3b scene IoU 0.953 -> 0.273; with compat 906 it scores 0.972, above the baseline, 0 degenerate. Root cause is upstream c7d8722922a setting prop.integrated on HIP builds, which triggers an MMQ tile-barrier race on gfx115x that silently corrupts output past n_ubatch (llama.cpp#28211). Upstream reverted it in #28604, 78 minutes after our payload was cut. Also corrects the layer-placement figures, which were from the gate matrix at a different context size, and records the hypotheses refuted along the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Answers the brief in
docs/maxusai/tasks/rocm-gate-issues-17459-17475.md, and commits theharness that produced it.
Results
ollama#17459 — 0
<unused49>tokens in 96 requests across eight payload/environmentcombinations (b9888; b10091 fork and stock; b10864 fork dio-on, fork dio-off, stock), cold
and warm,
thinkboth ways. On this hostthink=falseis the arm that works, which is theopposite of the report — on the same gfx1151 silicon the reporter used.
ollama#17475 — 0 contaminated victims in 436 extractions (484 including superseded runs)
across six configurations and all three protocols, with donor aborts and multi-model swapping
verified to have actually happened rather than assumed.
Three instrument corrections came first
Each produced a confident, wrong answer before it was found:
hardware; this host answers the same request in 3.9 s, so V2/V3 silently degraded into V1
with concurrency. Shortened to 2.5 s and confirmed server-side.
multi-model swapping, and with the reporter's 17–22 GB noise models the loop cycled 1–3
times across six configurations. Now runs until the loop has cycled, with a small-model
variant reaching 1203 cycles.
ZX9SYNTHETMARK448embedsSYNTHET, which alsoappears in the "SYNTHETIC TEST DOCUMENT" header on both forms — so victims that quoted their
own header scored as contaminated, 7 of 8 in one run, a false positive shaped almost exactly
like the reported bug. The generator now asserts against everything rendered on both forms.
The finding neither issue was about
0.32.x cannot see this box's VRAM. The scheduler reports 27 GiB where 0.34.x reports
95.4 GiB with the ROCm runtime held constant — stock
0.34.1-rocmships the same HIP7.2.70201 as stock
0.32.5-rocmand still sees 95.4. At an identical context,gemma4:31bplaces 25–26 of 61 layers on the GPU under 0.32.1 and 61 of 61 under 0.34.1. The
mechanism is not located and is reported as observed, not explained (ADR 0024).
Separately, direct I/O — which the gate treats as an unvalidated risk — is also what makes
repeated model loading viable here: 1203 completed noise cycles with it on against 3 with it
off, same image, same window.
What it means for the gate
Clauses 1 and 2 cannot be satisfied as written: both issues are closed
not_planned, so no fixcan ever be "in the target tag". Clause 3 now has an opt-out knob (
OLLAMA_IGPU_DIRECT_IO).Clause 4 — the vision A/B on the candidate — is the one remaining piece of work the gate
actually asks for, and PR #303 records the baseline it would be compared against.
The document does not recommend lifting the pin. It reports what reproduces, what does not,
what the pin costs, and leaves the decision where it belongs.
Caveats, stated in the document rather than buried
A clean run is not proof of absence for a concurrency race; our V3 deviates from the reporter's
in the noise models by necessity; the 0.32.x arms ran half on CPU because of the VRAM finding;
and ollama#17475's reporter is on aarch64 GB10 with the CUDA v13 runner, so a negative here does not
clear their configuration.
Harness
docs/maxusai/tasks/rocm-gate/— the reproduction scripts and the three table generators, sothe tables can be regenerated from the raw results rather than trusted (ADR 0012 rule 8). Host
store path and GPU GIDs are env-supplied, not committed. The README is explicit that the
harness binds production's model store read-write and why
OLLAMA_NOPRUNE=1is not optional.🤖 Generated with Claude Code