Pin llama.cpp 358cafc: ZAYA1-8B decodes 1.9x faster on HRX; concurrent ZAYA on Vulkan 2.5x - #168
Conversation
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
|
End-to-end through
|
Moves
third_party/llama.cppfrom8dd75ebto358cafc, the tip of1bit/hrx-vulkan-patched. This brings in two fork PRs:Measured on Strix Halo, ZAYA1-8B Q4_K_M,
GGML_HRX_DISABLE_DISPATCH=decode_splitasservesets:Correctness:
test-llama-archs -a zayapasses on HRX0, Vulkan and CPU.Other models on HRX0 (perplexity, 8 chunks, old build vs new):
GLM-4.7-Flash and Qwen3.6-35B-A3B weren't run: the GLM file isn't on the box, and the 35B Q8_0 is too big to load safely next to the other jobs right now. The weekly registry check runs both.
Also here:
registry_pins); counts unchanged.docs/vulkan.md: ZAYA's HRX row is 2,137 / 47.9 / 21.65, plus a paragraph on the change.docs/hrx.md: the new kernels are listed under Our patches.🤖 Generated with Claude Code