Pin llama.cpp a34a6b7: HRX decode-split multipass output pass is vectorised (engine#124) - #180
Conversation
…orised (engine#124) Brings 1bit-MONSTER/llama.cpp#30. The multipass reducer's output pass gave each workitem one channel and walked every KV block serially, recomputing expf(partial_max - maximum) per (block, element); only half the workgroup is live at value_head_size=128, which was the whole residual >2048 decode gap. The per-block scale is now computed once in the lane-strided sum pass into a per-row LDS stage, and one vectorised all-rows pass (4 channels per workitem) does the output. Per-channel block order is unchanged, so the reduce is bit-identical (1.24 GB of decode-path logits compare equal to the previous pin). d2100 +15.4%, d3000 +19.8%, d4800 +22.0%; d2100/d2000 0.881 -> 0.976, so the boundary cliff is gone. Nothing changes at or below capacity 2048. Details and repro in benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md.
|
Docs7 for 1bit-monster/engine
Commit |
85f8039 to
1e4847f
Compare
PR Reviewer Guide 🔍(Review updated until commit 20181e4)Here are some key observations to aid the review process:
|
|
Persistent review updated to latest commit 1e4847f |
|
Blocked on 1bit-MONSTER/llama.cpp#30: the new workgroup scale stage grows with context length (rows × blocks × 4 B) and can pass 64 KB of LDS at 64K-131K tokens, while the PR was measured only up to d4800. Details: 1bit-MONSTER/llama.cpp#30 (comment). Re-pin here once #30 bounds it or shows long-context runs are safe. |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Persistent review updated to latest commit 20181e4 |
…and PORTING status (#182) The decode-split q8 pack read global output behind an LDS-only barrier; fixed in llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised (#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Pins 1bit-MONSTER/llama.cpp#30 and documents it. Closes #124.
What
The multipass decode-split reducer (
reduce_completed.multipass, used abovekey_value_token_capacity2048) ended with a scalar output pass: each workitem owned one output channel and walked every KV block serially, recomputingexpf(partial_max - maximum)once per (block, element). Only half the workgroup is live atvalue_head_size = 128, so decode above 2048 stayed ~16% below the <=2048 corridor.The per-block scale is now computed once in the lane-strided sum pass into a per-row workgroup (LDS) stage, and one vectorised all-rows pass (4 channels per workitem) does the output. The per-channel block order is unchanged, so the reduce is bit-identical.
Measured (Strix Halo, Qwen3-Coder-30B-A3B-Instruct Q4_K_M,
-dev HRX0)llama-bench -p 0 -n 8 -r 5, median of 6 interleaved runs:d2100/d2000moves 0.881 -> 0.976 — the boundary cliff is gone. Nothing changes at or below capacity 2048.Evidence
llama-perplexity -c 2049 -b 1 --save-all-logits;GGML_HRX_LOG_DISPATCH=1confirmsflash_attention_decode_split) are byte-identical to the previous pin; PPL 21.9760 +/- 0.81813 on both.ZX-4718-QQexact at 4700 prompt tokens (capacity 4864, multipass), temperature 0, seed 42,cache_prompt:false.HSA_STATUS_ERROR_MEMORY_FAULT, 0res = -3.Full table, method and repro:
third_party/llama.cpp->benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md.Docs: new
docs/hrx.md"Fixed: decode-split multipass gap above 2048 (engine#124)" section.Pin updated: llama.cpp #30 is merged, so this now pins the branch head
fcd83eb(a merge commit with the same tree asa34a6b7);registry/architectures.jsonregenerated for it, andmainmerged in.Review (strixhalo, fork build of
a34a6b7vs its base00adc2b, same toolchain)barrier<workgroup>; the pack'sbarrier<global>+barrier<workgroup>from Pin llama.cpp 00adc2b: HRX decode-split publishes next_q8 after a global barrier (#123, #140) #176 is untouched.GGML_HRX_FA_PARTIAL_ALIGN=256, a separate process forcing KFD queue evictions: 8/8 correct, 0 NaN.llama-perplexity -c 2304 -b 1 --chunks 1, Coder-30B) = 7.5538 +/- 0.66265 on both builds.llama-bench -p 0 -n 16 -r 3, 2 interleaved rounds, tok/s): d2000 69.7/68.7 -> 67.5/67.3 (noise), d2100 48.9/49.7 -> 68.0/64.1, d4800 40.2/41.7 -> 42.4/47.0.scale_stagegrows with the block count. JIT-compiled at capacity 32768: 34.1 KB (32/4 heads) and 50.5 KB (64/4) vs a flat 17.7 KB before, under the 64 KB limit. Both builds fail to compile the kernel above 32768 (40960 to 262144 tried), so that ceiling is pre-existing and this PR does not move it.