From 1e4847fd86c55f0d81aa13df4800f1b7ba10f1e4 Mon Sep 17 00:00:00 2001 From: agent-872bf5 Date: Sun, 27 Sep 2026 18:51:31 -0300 Subject: [PATCH 1/2] Pin llama.cpp a34a6b7: HRX decode-split multipass output pass is vectorised (engine#124) Brings 1bit-MONSTER/llama.cpp#30. The multipass reducer's output pass gave each workitem one channel and walked every KV block serially, recomputing expf(partial_max - maximum) per (block, element); only half the workgroup is live at value_head_size=128, which was the whole residual >2048 decode gap. The per-block scale is now computed once in the lane-strided sum pass into a per-row LDS stage, and one vectorised all-rows pass (4 channels per workitem) does the output. Per-channel block order is unchanged, so the reduce is bit-identical (1.24 GB of decode-path logits compare equal to the previous pin). d2100 +15.4%, d3000 +19.8%, d4800 +22.0%; d2100/d2000 0.881 -> 0.976, so the boundary cliff is gone. Nothing changes at or below capacity 2048. Details and repro in benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md. --- docs/hrx.md | 28 ++++++++++++++++++++++++++++ third_party/llama.cpp | 2 +- 2 files changed, 29 insertions(+), 1 deletion(-) diff --git a/docs/hrx.md b/docs/hrx.md index 29dd80b..fe1b63f 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -117,6 +117,34 @@ Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan): `test-backend-ops -b HRX0`: 790/790. +### Fixed: decode-split multipass gap above 2048 ([engine#124](https://github.com/1bit-MONSTER/engine/issues/124)) + +The multipass decode-split reducer (`reduce_completed.multipass`, used above +`key_value_token_capacity` 2048) finished with a scalar output pass: each workitem owned +one output channel and walked every KV block serially, recomputing +`expf(partial_max - maximum)` once per (block, element). Only half the workgroup is live +at `value_head_size = 128`, so decode above 2048 stayed about 16% below the <=2048 +corridor. Fork PR +[1bit-MONSTER/llama.cpp#30](https://github.com/1bit-MONSTER/llama.cpp/pull/30) computes +the per-block scale once into a per-row workgroup (LDS) stage and runs one vectorised, +all-rows output pass (4 channels per workitem). The per-channel block order is +unchanged, so the reduce is bit-identical. + +Measured on Strix Halo (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, `-dev HRX0`, +`llama-bench -p 0 -n 8 -r 5`, median of 6 interleaved runs): + +| depth | blocks | before | this pin | delta | +|---|---:|---:|---:|---:| +| 2000 (<=2048 control) | 32 | 66.06 | 68.82 | +4.2% | +| 2100 | 33 | 58.19 | 67.15 | **+15.4%** | +| 3000 | 47 | 50.84 | 60.89 | **+19.8%** | +| 4800 | 76 | 41.83 | 51.02 | **+22.0%** | + +`d2100/d2000` moves 0.881 -> 0.976, so the boundary cliff is gone. 1.24 GB of +decode-path logits (`llama-perplexity -c 2049 -b 1 --save-all-logits`) are byte-identical +to the previous pin, the buried code word is exact at 4700 tokens, and no GPU faults +were observed. + ### Known issues on `HRX0` - **Several sequences per batch fail.** `llama-perplexity` with `n_seq` > 1 stops on an diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 00adc2b..a34a6b7 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 00adc2b9d6676f39cb3fcca8b5ad97621e653fae +Subproject commit a34a6b75d06cfd255365b0cecbadc2a95d0d4e54 From 20181e47168927fe74b5e363e7ecb1fe46d0cac8 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sun, 27 Sep 2026 19:14:09 -0300 Subject: [PATCH 2/2] Pin llama.cpp fcd83eb (llama.cpp #30 merged; same tree as a34a6b7) Co-Authored-By: Claude Opus 5.5 --- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index f38f0a2..f4053f5 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8", - "llama.cpp (hrx)": "00adc2b9d6676f39cb3fcca8b5ad97621e653fae", + "llama.cpp (hrx)": "fcd83eb2a4a86cdae5d65e114981470593a190aa", "zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb" }, "counts": { diff --git a/third_party/llama.cpp b/third_party/llama.cpp index a34a6b7..fcd83eb 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit a34a6b75d06cfd255365b0cecbadc2a95d0d4e54 +Subproject commit fcd83eb2a4a86cdae5d65e114981470593a190aa