Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions docs/hrx.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,8 @@ Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan):
| Qwen3-Coder-30B-A3B | faults 2 of 3 runs (80–88 tok/s when it survives) | 66–71 tok/s, 5 of 5 |
| ZAYA1-8B | 23.5 tok/s | 25.5 tok/s |

(ZAYA1-8B decodes at 47.9 tok/s since the HRX kernels below.)

A `GGML_HRX_DISABLE_DISPATCH` you set yourself wins, and `ONEBIT_HRX_DECODE_SPLIT=1` turns the
kernel back on, for testing a fix.

Expand Down Expand Up @@ -190,6 +192,13 @@ BF16 is 0.063.
views), which cannot move. Leaf (`NONE`) nodes count as covered. AMD's IQ4_NL
and IQ4_XS matmul kernels, and GET_ROWS for IQ4_XS and batched IQ3_S, give wrong
values, so those nodes are left to the CPU.
- **Kernels for ops HRX sent to the CPU** ([llama.cpp #24](https://github.com/1bit-MONSTER/llama.cpp/pull/24)):
a batched F16 matmul (`ggml_grouped_mul_mat_f16_f32`, e.g. ZAYA's grouped convolution, whose
weights the loader can now place on HRX), and short-row kernels in `small_rows_f32.loom`:
SOFT_MAX without a mask, SUM_ROWS, ARGSORT, GET_ROWS for narrow rows (strided ids, as top-k views
are), CONT of strided views and broadcast-only REPEAT. Each matcher claims only what the existing
kernels do not. `test-backend-ops -b HRX0` passes every case of these ops. ZAYA1-8B's decode graph
goes from 641 graph splits to 1: 25.5 to 47.9 tok/s (tg128).

Measured on Strix Halo (Qwen3-0.6B, perplexity over 8 x 512 wikitext tokens):

Expand Down
9 changes: 8 additions & 1 deletion docs/vulkan.md
Original file line number Diff line number Diff line change
Expand Up @@ -172,7 +172,7 @@ Q4_K_M (5.17 GiB), by device:
|---|---|---|---|
| Vulkan0 | 3,437 tok/s | 93.0 tok/s | 21.57 |
| ROCm0 | ~2,400 tok/s | 61.5 tok/s | 21.78 |
| HRX0 | 1,175 tok/s | 25.5 tok/s | 21.55 |
| HRX0 | 2,137 tok/s | 47.9 tok/s | 21.65 |

F16 on Vulkan0: 1,138 tok/s prefill, 45.9 tok/s decode, perplexity 20.59.

Expand All @@ -184,6 +184,13 @@ Vulkan decodes fastest, so `auto` and `--device vulkan` stay the default for ZAY
`--device rocm` server is built from ROCmFPX's tree, which has no ZAYA; the ROCm numbers above are
our llama.cpp built with `GGML_HIP=ON` for gfx1151.

HRX0 decoded at 25.5 tok/s until [llama.cpp #24](https://github.com/1bit-MONSTER/llama.cpp/pull/24): ggml-hrx sent
seven of ZAYA's ops per layer to the CPU (the grouped-conv matmul, the router's softmax, top-k and
gather, and a few copies), 641 graph splits per decoded token. New HRX kernels for those ops (see
[hrx.md](hrx.md#our-patches)) make the decode graph one split, and decode 1.9x faster. The perplexity moves
from 21.55 to 21.65, which is kernel rounding amplified by top-1 routing: with the new kernels turned off
(`GGML_HRX_DISABLE_DISPATCH`) the build reproduces the old figure exactly. Vulkan and the CPU are unchanged.

What it took, besides the port: CCA's grouped convolution runs as one batched matmul per tap
(as one small matmul per group it held Vulkan decode at 50 tok/s), and the graph avoids what
HRX and ROCm lacked: copies into part of a state row, concatenating strided views, `l2_norm`,
Expand Down
2 changes: 1 addition & 1 deletion registry/architectures.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
"llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8",
"llama.cpp (hrx)": "8dd75eb48f282198abdb0bd748aa0e4594745f1d",
"llama.cpp (hrx)": "358cafc249a0234e5c7ffc46201f20aa6682b452",
"zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb"
},
"counts": {
Expand Down
Loading