From 433f7a3f21a1fb86f278502e22c3714dcdc32b4b Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sat, 26 Sep 2026 20:25:50 -0300 Subject: [PATCH] Pin llama.cpp 358cafc: HRX kernels for ZAYA's router and grouped conv, ZAYA1-8B decode 25.5 -> 47.9 tok/s on HRX (llama.cpp #22, #24) Co-Authored-By: Claude Opus 5.5 --- docs/hrx.md | 9 +++++++++ docs/vulkan.md | 9 ++++++++- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 4 files changed, 19 insertions(+), 3 deletions(-) diff --git a/docs/hrx.md b/docs/hrx.md index 431f7b8..aaed81b 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -134,6 +134,8 @@ Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan): | Qwen3-Coder-30B-A3B | faults 2 of 3 runs (80–88 tok/s when it survives) | 66–71 tok/s, 5 of 5 | | ZAYA1-8B | 23.5 tok/s | 25.5 tok/s | + (ZAYA1-8B decodes at 47.9 tok/s since the HRX kernels below.) + A `GGML_HRX_DISABLE_DISPATCH` you set yourself wins, and `ONEBIT_HRX_DECODE_SPLIT=1` turns the kernel back on, for testing a fix. @@ -190,6 +192,13 @@ BF16 is 0.063. views), which cannot move. Leaf (`NONE`) nodes count as covered. AMD's IQ4_NL and IQ4_XS matmul kernels, and GET_ROWS for IQ4_XS and batched IQ3_S, give wrong values, so those nodes are left to the CPU. +- **Kernels for ops HRX sent to the CPU** ([llama.cpp #24](https://github.com/1bit-MONSTER/llama.cpp/pull/24)): + a batched F16 matmul (`ggml_grouped_mul_mat_f16_f32`, e.g. ZAYA's grouped convolution, whose + weights the loader can now place on HRX), and short-row kernels in `small_rows_f32.loom`: + SOFT_MAX without a mask, SUM_ROWS, ARGSORT, GET_ROWS for narrow rows (strided ids, as top-k views + are), CONT of strided views and broadcast-only REPEAT. Each matcher claims only what the existing + kernels do not. `test-backend-ops -b HRX0` passes every case of these ops. ZAYA1-8B's decode graph + goes from 641 graph splits to 1: 25.5 to 47.9 tok/s (tg128). Measured on Strix Halo (Qwen3-0.6B, perplexity over 8 x 512 wikitext tokens): diff --git a/docs/vulkan.md b/docs/vulkan.md index 9f8fc25..9c02e82 100644 --- a/docs/vulkan.md +++ b/docs/vulkan.md @@ -172,7 +172,7 @@ Q4_K_M (5.17 GiB), by device: |---|---|---|---| | Vulkan0 | 3,437 tok/s | 93.0 tok/s | 21.57 | | ROCm0 | ~2,400 tok/s | 61.5 tok/s | 21.78 | -| HRX0 | 1,175 tok/s | 25.5 tok/s | 21.55 | +| HRX0 | 2,137 tok/s | 47.9 tok/s | 21.65 | F16 on Vulkan0: 1,138 tok/s prefill, 45.9 tok/s decode, perplexity 20.59. @@ -184,6 +184,13 @@ Vulkan decodes fastest, so `auto` and `--device vulkan` stay the default for ZAY `--device rocm` server is built from ROCmFPX's tree, which has no ZAYA; the ROCm numbers above are our llama.cpp built with `GGML_HIP=ON` for gfx1151. +HRX0 decoded at 25.5 tok/s until [llama.cpp #24](https://github.com/1bit-MONSTER/llama.cpp/pull/24): ggml-hrx sent +seven of ZAYA's ops per layer to the CPU (the grouped-conv matmul, the router's softmax, top-k and +gather, and a few copies), 641 graph splits per decoded token. New HRX kernels for those ops (see +[hrx.md](hrx.md#our-patches)) make the decode graph one split, and decode 1.9x faster. The perplexity moves +from 21.55 to 21.65, which is kernel rounding amplified by top-1 routing: with the new kernels turned off +(`GGML_HRX_DISABLE_DISPATCH`) the build reproduces the old figure exactly. Vulkan and the CPU are unchanged. + What it took, besides the port: CCA's grouped convolution runs as one batched matmul per tap (as one small matmul per group it held Vulkan decode at 50 tok/s), and the graph avoids what HRX and ROCm lacked: copies into part of a state row, concatenating strided views, `l2_norm`, diff --git a/registry/architectures.json b/registry/architectures.json index aa49ede..9ad3d87 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8", - "llama.cpp (hrx)": "8dd75eb48f282198abdb0bd748aa0e4594745f1d", + "llama.cpp (hrx)": "358cafc249a0234e5c7ffc46201f20aa6682b452", "zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb" }, "counts": { diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 8dd75eb..358cafc 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 8dd75eb48f282198abdb0bd748aa0e4594745f1d +Subproject commit 358cafc249a0234e5c7ffc46201f20aa6682b452