Skip to content

Pin llama.cpp 358cafc: ZAYA1-8B decodes 1.9x faster on HRX; concurrent ZAYA on Vulkan 2.5x - #168

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin/hrx-zaya-kernels
Sep 26, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin/hrx-zaya-kernels

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from 8dd75eb to 358cafc, the tip of 1bit/hrx-vulkan-patched. This brings in two fork PRs:

  • llama.cpp #22: ZAYA's grouped conv folds the sequences into the token dimension. Vulkan can't broadcast a matmul weight over dim 3, so whenever a ubatch could hold several sequences both tap matmuls of every layer ran on the CPU (162 splits). With 4 concurrent requests on Vulkan0, total decode goes from 50–66 to 135–160 tok/s. One request is unchanged.
  • llama.cpp #24: new HRX kernels for the ops ZAYA sent to the CPU (batched F16 matmul, row softmax/sum/argsort, narrow gather, strided copy, broadcast repeat), plus small ZAYA graph changes. The decode graph on HRX0 goes from 641 splits per token to 1.

Measured on Strix Halo, ZAYA1-8B Q4_K_M, GGML_HRX_DISABLE_DISPATCH=decode_split as serve sets:

before after
HRX0 decode, tg128 25.5 tok/s 47.9 tok/s
HRX0 prefill, pp512 1,175 tok/s 2,137 tok/s
Vulkan0 decode (server) 77–83 tok/s 83–84 tok/s

Correctness:

  • Wikitext perplexity is unchanged on Vulkan0 (21.5731, and 21.6692 with 4 sequences) and on the CPU (25.1372 with 1 and 4 sequences).
  • HRX0 goes 21.5518 → 21.6456. That's rounding in the new kernels, amplified by top-1 routing: with them turned off the build reproduces the old figure exactly.
  • Teacher-forced top-1 against transformers FP32 is unchanged: HRX0 69/96 and Vulkan0 70/96 for Q4_K_M, and F16 95/96.
  • test-llama-archs -a zaya passes on HRX0, Vulkan and CPU.

Other models on HRX0 (perplexity, 8 chunks, old build vs new):

Model old new
Qwen3-0.6B 39.3098 39.3098
Qwen3-Coder-30B-A3B 17.2526 17.2526
Qwen2.5-7B 18.7171 18.7171
MiniCPM4-8B 20.2310 20.2310

GLM-4.7-Flash and Qwen3.6-35B-A3B weren't run: the GLM file isn't on the box, and the 35B Q8_0 is too big to load safely next to the other jobs right now. The weekly registry check runs both.

Also here:

  • The registry is regenerated for the pin (registry_pins); counts unchanged.
  • docs/vulkan.md: ZAYA's HRX row is 2,137 / 47.9 / 21.65, plus a paragraph on the change.
  • docs/hrx.md: the new kernels are listed under Our patches.

🤖 Generated with Claude Code

…, ZAYA1-8B decode 25.5 -> 47.9 tok/s on HRX (llama.cpp #22, #24)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 26, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 433f7a3

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

22 - Partially compliant

Compliant requirements:

  • Removed third_party/lemonade and related Lemonade-specific tests
  • Moved device support into 1bit serve
  • Pinned dependencies in the engine
  • Updated portable child handling
  • Added tests/smoke_serve.sh and tests/fake_backend.py
  • Updated documentation to reflect new relationship
  • Verified builds and tests on Ryzen, Strix Halo NPU, and Vulkan

Non-compliant requirements:

  • Not rerun --device mlx on Mac (but launch contract is unchanged)

Requires further human verification:

  • Verification of --device mlx on Mac

24 - Partially compliant

Compliant requirements:

  • Updated docs/lemonade.md
  • Updated README status and PORTING step 1

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Documentation Inconsistency

The documentation mentions that ZAYA1-8B decodes at 47.9 tok/s due to HRX kernels, but this performance improvement is only measured on Strix Halo. The documentation should clarify that this performance gain is specific to the Strix Halo hardware and may not be reproducible on other hardware configurations.

(ZAYA1-8B decodes at 47.9 tok/s since the HRX kernels below.)
Documentation Inconsistency

The documentation states that HRX0 decoded at 25.5 tok/s until llama.cpp #24, but then provides a comparison with Vulkan0 decode performance. It's unclear if this comparison is meant to show relative performance or absolute performance, and the documentation should clarify this.

HRX0 decoded at 25.5 tok/s until [llama.cpp #24](https://github.com/1bit-MONSTER/llama.cpp/pull/24): ggml-hrx sent
seven of ZAYA's ops per layer to the CPU (the grouped-conv matmul, the router's softmax, top-k and
gather, and a few copies), 641 graph splits per decoded token. New HRX kernels for those ops (see
[hrx.md](hrx.md#our-patches)) make the decode graph one split, and decode 1.9x faster. The perplexity moves
from 21.55 to 21.65, which is kernel rounding amplified by top-1 routing: with the new kernels turned off
(`GGML_HRX_DISABLE_DISPATCH`) the build reproduces the old figure exactly. Vulkan and the CPU are unchanged.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

End-to-end through 1bit serve on strixhalo. This PR's engine was built against the llama-server from the same tree as 358cafc.

  • serve_e2e ZAYA1-8B Q4_K_M: PASS on --device hrx and --device vulkan.
  • serve_e2e Qwen3-Coder-30B-A3B --device hrx: 5 of 6 runs pass (the last 4 in a row). The first run failed and its output wasn't kept; the box had 23 GB free at the time. Coder-30B's perplexity is identical to before (17.2526), so it's not a numerics change. If the weekly check sees it again, it deserves its own issue.

@bong-water-water-bong
bong-water-water-bong merged commit cfb0b46 into main Sep 26, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin/hrx-zaya-kernels branch September 26, 2026 23:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant