preflight: gfx1151's output-quality floors and fp16 canary, measured on production's configuration - #403
Conversation
|
Merge order with #402. Both PRs regenerate the README's release matrix. A test merge of #403 on top of #402 conflicts only in the matrix's source comment (#402 lists the Metal production run, #403 the gfx1151 baseline run). The table rows combine cleanly. Suggested order: #402 first. I will then rebase #403, regenerate the table from all three hosts' runs, and merge both comments. #400 merges cleanly with either.
|
…on production's configuration
The rocm7-0-34-4-dynres profile recorded no quality thresholds and no poison-probe expectation, so its
"Output quality" and "fp16 overflow canary" checks read SKIP on every gfx1151 run. This records both, from
2026-09-28's baseline of the promoted image (0.34.4-dynres-0-gb43ee8e) in a bench container with production's
environment: f16, flash attention on, two-pass, two slots.
- [quality.rocm7-0-34-4-dynres.{nemotron_h_omni,gemma4,qwen35}]: the three arches' models measured 6/6 and 5/5
with valid JSON, think off; the floors are cuda-dynres-005's (1.0, 0.70, 0.70).
- [poison.rocm7-0-34-4-dynres]: the canary's first run on HIP. It passes under the fork's f32 gate and fails
under GGML_CUDA_CUBLAS_COMPUTE_TYPE=f16 ('?' until num_predict; 6 inf at ffn_down-31), so it tells the gate
apart here as on CUDA.
- The preflight with both: VERDICT PASS, PASS=24 SKIP=8 (was 20/12). Its run record is committed, and the
README's generated matrix now takes gfx1151's row from it.
- amd-upgrade-gate.md records the baseline: think off on five models under f16 against the q8_0 gate cells
(all 135 blocks valid; 37-72 cells move per model, mostly toward f16), and OCRBench 172/200 = q8_0's, on the
same items.
- test_verdicts.py now reaches SUITE_DIR, as check_quality does. CI runs it from preflight/, where "." does not
reach summarize_engine_compare, so TestQualityThresholds skipped on the full tree (6 tests); they now run and
pass. A preflight-only tree still skips them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
6ec965d to
b38f168
Compare
|
Rebased onto main after #402, in b38f168. The README matrix is regenerated from all three hosts' runs, and it equals
|
The
rocm7-0-34-4-dynresprofile recorded no quality thresholds and no poison-probe expectation. So on every gfx1151 run, its "Output quality" and "fp16 overflow canary" checks read SKIP. This records both from a new baseline of what production runs, and records the baseline indocs/maxusai/amd-upgrade-gate.md.The baseline (2026-09-28, at the maintainer's word)
The promoted image,
0.34.4-dynres-0-gb43ee8e, ran in a bench container on:11497with production's environment: f16, flash attention on, two-pass and two slots. Production was not touched. The server logs show f16 K and V, flash attention and two-pass on every load. Every model ran at its automatic batch, 2048 for gemma4 and 1024 for the others, and no images were decoded in pieces.q8_0cells, the KV type moves 37 to 72 cells per model. More scored moves favour f16 on four models; qwen3.8 is close to even, 8 against 9.q8_0, on the same items (McNemar p = 1.000). 199 of 200 answers are the same text.What it records
[quality.rocm7-0-34-4-dynres.{nemotron_h_omni,gemma4,qwen35}]. The three arches' models scored 6/6 onscene_singleand 5/5 ondocument_single, with valid JSON. The floors arecuda-dynres-005's: JSON 1.0, label recall 0.70 and qty/price 0.70, which allow one miss per test.[poison.rocm7-0-34-4-dynres], the canary's first run on HIP. Under the fork's f32 gate the trigger decodes correctly, and the node meter is clean over 33 nodes. WithGGML_CUDA_CUBLAS_COMPUTE_TYPE=f16set on the container, the decode is?untilnum_predict.ffn_down-31then holds 6 infinities, and the next metered node holds 12,288 NaNs. So the check tells the gate apart on gfx1151, as on CUDA. The entry went in only because it did: a canary that cannot fail here would put a false green in the matrix.A CI gap this found
CI runs
test_verdicts.pyfrompreflight/. From there,"."does not reach the suite'ssummarize_engine_compare, so all sixTestQualityThresholdstests skipped on the full tree. A recent run showsOK (skipped=6). The file now appendschecks.SUITE_DIRtosys.path, ascheck_qualitydoes at runtime. All 196 tests run and pass, and a preflight-only tree still skips the six, as CI's release-lineage step expects.Still running
Think-on runs overnight under f16 and two-pass: gemma4:31b and gemma4:26b greedy at a pinned batch of 2048, and qwen3.8 and nemotron3 twice each at their packaged sampling. The result will be added to this PR when it ends. qwen3.6 already has this configuration (
fold2punder f16 in the KV task record).Verification
test_verdicts.py: 196 tests, OK.check_source_paths.py --changed-since origin/mainis clean.release_matrix.py --version 0.34.4-dynres runs/*.json, byte for byte.amd-server/rocm-gfx1151🤖 Generated with Claude Code