Skip to content

docs: ROCm gate — #17459 and #17475 do not reproduce on gfx1151 - #304

Merged
glennneuber merged 2 commits into
mainfrom
docs/rocm-gate-issues-result
Sep 18, 2026
Merged

glennneuber merged 2 commits into
mainfrom
docs/rocm-gate-issues-result

Conversation

@glennneuber

Copy link
Copy Markdown

Answers the brief in docs/maxusai/tasks/rocm-gate-issues-17459-17475.md, and commits the
harness that produced it.

Results

ollama#17459 — 0 <unused49> tokens in 96 requests across eight payload/environment
combinations (b9888; b10091 fork and stock; b10864 fork dio-on, fork dio-off, stock), cold
and warm, think both ways. On this host think=false is the arm that works, which is the
opposite of the report — on the same gfx1151 silicon the reporter used.

ollama#17475 — 0 contaminated victims in 436 extractions (484 including superseded runs)
across six configurations and all three protocols, with donor aborts and multi-model swapping
verified to have actually happened rather than assumed.

Three instrument corrections came first

Each produced a confident, wrong answer before it was found:

  1. The abort never fired. The reporter's 8 s timeout lands mid-generation on their
    hardware; this host answers the same request in 3.9 s, so V2/V3 silently degraded into V1
    with concurrency. Shortened to 2.5 s and confirmed server-side.
  2. The swapping never happened. V3's only distinguishing ingredient is continuous
    multi-model swapping, and with the reporter's 17–22 GB noise models the loop cycled 1–3
    times across six configurations. Now runs until the loop has cycled, with a small-model
    variant reaching 1203 cycles.
  3. The marker collided with the documents. ZX9SYNTHETMARK448 embeds SYNTHET, which also
    appears in the "SYNTHETIC TEST DOCUMENT" header on both forms — so victims that quoted their
    own header scored as contaminated, 7 of 8 in one run, a false positive shaped almost exactly
    like the reported bug. The generator now asserts against everything rendered on both forms.

The finding neither issue was about

0.32.x cannot see this box's VRAM. The scheduler reports 27 GiB where 0.34.x reports
95.4 GiB with the ROCm runtime held constant — stock 0.34.1-rocm ships the same HIP
7.2.70201 as stock 0.32.5-rocm and still sees 95.4. At an identical context, gemma4:31b
places 25–26 of 61 layers on the GPU under 0.32.1 and 61 of 61 under 0.34.1. The
mechanism is not located and is reported as observed, not explained (ADR 0024).

Separately, direct I/O — which the gate treats as an unvalidated risk — is also what makes
repeated model loading viable here: 1203 completed noise cycles with it on against 3 with it
off, same image, same window.

What it means for the gate

Clauses 1 and 2 cannot be satisfied as written: both issues are closed not_planned, so no fix
can ever be "in the target tag". Clause 3 now has an opt-out knob (OLLAMA_IGPU_DIRECT_IO).
Clause 4 — the vision A/B on the candidate — is the one remaining piece of work the gate
actually asks for
, and PR #303 records the baseline it would be compared against.

The document does not recommend lifting the pin. It reports what reproduces, what does not,
what the pin costs, and leaves the decision where it belongs.

Caveats, stated in the document rather than buried

A clean run is not proof of absence for a concurrency race; our V3 deviates from the reporter's
in the noise models by necessity; the 0.32.x arms ran half on CPU because of the VRAM finding;
and ollama#17475's reporter is on aarch64 GB10 with the CUDA v13 runner, so a negative here does not
clear their configuration.

Harness

docs/maxusai/tasks/rocm-gate/ — the reproduction scripts and the three table generators, so
the tables can be regenerated from the raw results rather than trusted (ADR 0012 rule 8). Host
store path and GPU GIDs are env-supplied, not committed. The README is explicit that the
harness binds production's model store read-write and why OLLAMA_NOPRUNE=1 is not optional.

🤖 Generated with Claude Code

…fx1151

Answers the brief in docs/maxusai/tasks/rocm-gate-issues-17459-17475.md.

ollama#17459: 0 <unused49> tokens in 96 requests across eight payload/environment
combinations (b9888; b10091 fork and stock; b10864 fork dio-on, fork dio-off,
stock), cold and warm, think both ways. On this host think=false is the arm
that works, the opposite of the report.

ollama#17475: 0 contaminated victims in 436 extractions (484 including superseded
runs) across six configurations and all three protocols, with donor aborts and
multi-model swapping verified to have actually occurred. Three instrument
corrections were needed first: the 2.5 s abort window (8 s never fired here),
continuing victims until the noise loop cycled (V3 had swapped 1-3 times), and
a marker that does not collide with "SYNTHETIC TEST DOCUMENT" on the forms.

Incidental and more consequential than either issue: 0.32.x reports 27 GiB of
iGPU memory where 0.34.x reports 95.4 GiB with the ROCm runtime held constant,
so production places 25-26 of gemma4:31b's 61 layers on the GPU where 0.34.1
places 61. Observed, not explained. Direct I/O, which the gate treats as a
risk, is also what makes repeated model loading viable here.

Neither issue can satisfy its gate clause as written -- both are closed
not_planned, so no fix can be "in the target tag". Clause 4 (vision A/B on the
candidate) is the remaining piece of work the gate actually asks for.

Harness committed under docs/maxusai/tasks/rocm-gate/ so the reproduction can
be repeated or disputed; host store path and GPU GIDs are env-supplied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

Update that changes this PR's conclusion: clause 4 has now been run, and the candidate fails it.

This PR argues the pin rests on two unreproducible bugs plus a measured cost, and names the vision A/B on qwen35moe as "the one remaining piece of work the gate actually asks for". That work is done. The result is not the one this PR anticipated.

Clause 4: FAIL

Full vision campaign, five models, 0.34.1-dynres-649fea9e against the 2026-09-17 baseline (PR #303), same runner, same digests, same env, one power envelope:

model scene IoU 0.32.1 → 0.34.1 objects
qwen3.8:27b-q4_K_M 0.991 → 0.056 6/6 → 3/6
gemma4:31b-it-q4_K_M 0.961 → 0.922 6/6 → 3/6
nemotron3:33b-q4_K_M 0.857 → 0.161 6/6 → 3/6
qwen3.6:35b-a3b-q4_k_m (the gate model) 0.953 → 0.273 6/6 =
gemma4:26b-a4b-it-q4_K_M 0.973 → 0.966 6/6 =

Fine-text OCR collapses on the same three (4/4/4 → 0/0/0 on qwen3.8), invoice extraction 5/5 → 0/5. qwen35moe — the model the gate was tripped on — drops scene IoU 0.953 → 0.273 and loses the multi-image question. The pin does not lift on this evidence.

What that does to this document

The framing needs revising, and I would rather say so here than let the PR merge as written:

The regression, narrowed

Subsequent work localized it considerably — details in the comment on #309. Briefly: with the image-token budget pinned per cell, gemma4:31b name_bbox_mean_iou is 0.728 → 0.000 at the 1120 rung only; 280 and 560 are equivalent on both builds. prompt_eval_count is identical between builds, so the grid and token count are unchanged and only pixel values differ. It reproduces on stock ollama/ollama:0.34.1-rocm, so it is upstream rather than our compat series.

Controls held: the pinned build re-run today reproduces its own baseline exactly (host has not drifted); unaffected by direct I/O; unaffected by layer placement.

Mechanism still unknown. Refuted by experiment: fp16 accumulation in the vision tower, multi-sub-batch image decode, and reverting the b10864 resize-algo flip.

Suggested disposition

I would rather not merge this as-is. Either it gets revised to carry the clause 4 result, or it merges as a dated record of what was known on 2026-09-17 with a pointer to the regression work — but it should not stand as the current view of whether the gate can lift.

🤖 Generated with Claude Code

…906 fixes it

Clause 4 was run 2026-09-18. Unpatched b10864 dropped qwen3.6:35b-a3b scene IoU
0.953 -> 0.273; with compat 906 it scores 0.972, above the baseline, 0 degenerate.

Root cause is upstream c7d8722922a setting prop.integrated on HIP builds, which
triggers an MMQ tile-barrier race on gfx115x that silently corrupts output past
n_ubatch (llama.cpp#28211). Upstream reverted it in #28604, 78 minutes after our
payload was cut.

Also corrects the layer-placement figures, which were from the gate matrix at a
different context size, and records the hypotheses refuted along the way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit ba37238 into main Sep 18, 2026
1 check passed
@glennneuber
glennneuber deleted the docs/rocm-gate-issues-result branch September 18, 2026 11:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant