Skip to content

preflight: gfx1151's output-quality floors and fp16 canary, measured on production's configuration - #403

Merged
glennneuber merged 1 commit into
mainfrom
preflight/rocm7-quality-floors
Sep 28, 2026
Merged

glennneuber merged 1 commit into
mainfrom
preflight/rocm7-quality-floors

Conversation

@glennneuber

Copy link
Copy Markdown

The rocm7-0-34-4-dynres profile recorded no quality thresholds and no poison-probe expectation. So on every gfx1151 run, its "Output quality" and "fp16 overflow canary" checks read SKIP. This records both from a new baseline of what production runs, and records the baseline in docs/maxusai/amd-upgrade-gate.md.

The baseline (2026-09-28, at the maintainer's word)

The promoted image, 0.34.4-dynres-0-gb43ee8e, ran in a bench container on :11497 with production's environment: f16, flash attention on, two-pass and two slots. Production was not touched. The server logs show f16 K and V, flash attention and two-pass on every load. Every model ran at its automatic batch, 2048 for gemma4 and 1024 for the others, and no images were decoded in pieces.

  • Think off, five models: all 135 blocks finished with valid JSON. Against the gate's q8_0 cells, the KV type moves 37 to 72 cells per model. More scored moves favour f16 on four models; qwen3.8 is close to even, 8 against 9.
  • OCRBench, gemma4:31b, rows 0–200: 172/200 under f16, as under q8_0, on the same items (McNemar p = 1.000). 199 of 200 answers are the same text.

What it records

  • [quality.rocm7-0-34-4-dynres.{nemotron_h_omni,gemma4,qwen35}]. The three arches' models scored 6/6 on scene_single and 5/5 on document_single, with valid JSON. The floors are cuda-dynres-005's: JSON 1.0, label recall 0.70 and qty/price 0.70, which allow one miss per test.
  • [poison.rocm7-0-34-4-dynres], the canary's first run on HIP. Under the fork's f32 gate the trigger decodes correctly, and the node meter is clean over 33 nodes. With GGML_CUDA_CUBLAS_COMPUTE_TYPE=f16 set on the container, the decode is ? until num_predict. ffn_down-31 then holds 6 infinities, and the next metered node holds 12,288 NaNs. So the check tells the gate apart on gfx1151, as on CUDA. The entry went in only because it did: a canary that cannot fail here would put a false green in the matrix.
  • The preflight with both entries: VERDICT PASS, PASS=24 SKIP=8, against 20 and 12 before. The four new checks pass, and the eight remaining skips are checks that do not apply here or that no profile records (the aspect ladder). The run record is committed. The README's generated matrix now takes gfx1151's row from it, and the row is green in every column. The README's comment says the run came from a bench container, not from the production container.

A CI gap this found

CI runs test_verdicts.py from preflight/. From there, "." does not reach the suite's summarize_engine_compare, so all six TestQualityThresholds tests skipped on the full tree. A recent run shows OK (skipped=6). The file now appends checks.SUITE_DIR to sys.path, as check_quality does at runtime. All 196 tests run and pass, and a preflight-only tree still skips the six, as CI's release-lineage step expects.

Still running

Think-on runs overnight under f16 and two-pass: gemma4:31b and gemma4:26b greedy at a pinned batch of 2048, and qwen3.8 and nemotron3 twice each at their packaged sampling. The result will be added to this PR when it ends. qwen3.6 already has this configuration (fold2p under f16 in the KV task record).

Verification

  • test_verdicts.py: 196 tests, OK.
  • check_source_paths.py --changed-since origin/main is clean.
  • The README table equals release_matrix.py --version 0.34.4-dynres runs/*.json, byte for byte.
  • The run record contains no host paths or names.

amd-server/rocm-gfx1151

🤖 Generated with Claude Code

@glennneuber

Copy link
Copy Markdown
Author

Merge order with #402. Both PRs regenerate the README's release matrix. A test merge of #403 on top of #402 conflicts only in the matrix's source comment (#402 lists the Metal production run, #403 the gfx1151 baseline run). The table rows combine cleanly. Suggested order: #402 first. I will then rebase #403, regenerate the table from all three hosts' runs, and merge both comments. #400 merges cleanly with either.

amd-server/rocm-gfx1151

…on production's configuration

The rocm7-0-34-4-dynres profile recorded no quality thresholds and no poison-probe expectation, so its
"Output quality" and "fp16 overflow canary" checks read SKIP on every gfx1151 run. This records both, from
2026-09-28's baseline of the promoted image (0.34.4-dynres-0-gb43ee8e) in a bench container with production's
environment: f16, flash attention on, two-pass, two slots.

- [quality.rocm7-0-34-4-dynres.{nemotron_h_omni,gemma4,qwen35}]: the three arches' models measured 6/6 and 5/5
  with valid JSON, think off; the floors are cuda-dynres-005's (1.0, 0.70, 0.70).
- [poison.rocm7-0-34-4-dynres]: the canary's first run on HIP. It passes under the fork's f32 gate and fails
  under GGML_CUDA_CUBLAS_COMPUTE_TYPE=f16 ('?' until num_predict; 6 inf at ffn_down-31), so it tells the gate
  apart here as on CUDA.
- The preflight with both: VERDICT PASS, PASS=24 SKIP=8 (was 20/12). Its run record is committed, and the
  README's generated matrix now takes gfx1151's row from it.
- amd-upgrade-gate.md records the baseline: think off on five models under f16 against the q8_0 gate cells
  (all 135 blocks valid; 37-72 cells move per model, mostly toward f16), and OCRBench 172/200 = q8_0's, on the
  same items.
- test_verdicts.py now reaches SUITE_DIR, as check_quality does. CI runs it from preflight/, where "." does not
  reach summarize_engine_compare, so TestQualityThresholds skipped on the full tree (6 tests); they now run and
  pass. A preflight-only tree still skips them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber force-pushed the preflight/rocm7-quality-floors branch from 6ec965d to b38f168 Compare September 28, 2026 06:36
@glennneuber

Copy link
Copy Markdown
Author

Rebased onto main after #402, in b38f168. The README matrix is regenerated from all three hosts' runs, and it equals release_matrix.py --version 0.34.4-dynres runs/*.json byte for byte. Its source comment keeps #402's list and names the gfx1151 baseline run as the newest for rocm7. test_verdicts.py passes (OK), and the source-path check is clean.

amd-server/rocm-gfx1151

@glennneuber
glennneuber merged commit 67499b2 into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant