diff --git a/docs/hrx.md b/docs/hrx.md index 29dd80b..fe1b63f 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -117,6 +117,34 @@ Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan): `test-backend-ops -b HRX0`: 790/790. +### Fixed: decode-split multipass gap above 2048 ([engine#124](https://github.com/1bit-MONSTER/engine/issues/124)) + +The multipass decode-split reducer (`reduce_completed.multipass`, used above +`key_value_token_capacity` 2048) finished with a scalar output pass: each workitem owned +one output channel and walked every KV block serially, recomputing +`expf(partial_max - maximum)` once per (block, element). Only half the workgroup is live +at `value_head_size = 128`, so decode above 2048 stayed about 16% below the <=2048 +corridor. Fork PR +[1bit-MONSTER/llama.cpp#30](https://github.com/1bit-MONSTER/llama.cpp/pull/30) computes +the per-block scale once into a per-row workgroup (LDS) stage and runs one vectorised, +all-rows output pass (4 channels per workitem). The per-channel block order is +unchanged, so the reduce is bit-identical. + +Measured on Strix Halo (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, `-dev HRX0`, +`llama-bench -p 0 -n 8 -r 5`, median of 6 interleaved runs): + +| depth | blocks | before | this pin | delta | +|---|---:|---:|---:|---:| +| 2000 (<=2048 control) | 32 | 66.06 | 68.82 | +4.2% | +| 2100 | 33 | 58.19 | 67.15 | **+15.4%** | +| 3000 | 47 | 50.84 | 60.89 | **+19.8%** | +| 4800 | 76 | 41.83 | 51.02 | **+22.0%** | + +`d2100/d2000` moves 0.881 -> 0.976, so the boundary cliff is gone. 1.24 GB of +decode-path logits (`llama-perplexity -c 2049 -b 1 --save-all-logits`) are byte-identical +to the previous pin, the buried code word is exact at 4700 tokens, and no GPU faults +were observed. + ### Known issues on `HRX0` - **Several sequences per batch fail.** `llama-perplexity` with `n_seq` > 1 stops on an diff --git a/registry/architectures.json b/registry/architectures.json index f38f0a2..f4053f5 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8", - "llama.cpp (hrx)": "00adc2b9d6676f39cb3fcca8b5ad97621e653fae", + "llama.cpp (hrx)": "fcd83eb2a4a86cdae5d65e114981470593a190aa", "zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb" }, "counts": { diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 00adc2b..fcd83eb 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 00adc2b9d6676f39cb3fcca8b5ad97621e653fae +Subproject commit fcd83eb2a4a86cdae5d65e114981470593a190aa