From 7c3720185065db33370b01061a8b4afd30062aeb Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sun, 27 Sep 2026 11:45:43 -0300 Subject: [PATCH] Pin llama.cpp fa226f9: ZAYA chat templates tokenize (llama.cpp #27); HRX makes an all-NaN MoE router row loud (llama.cpp #26, engine#123) Co-Authored-By: Claude Opus 5.5 --- docs/hrx.md | 11 +++++++++-- docs/vulkan.md | 8 ++++++++ registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 4 files changed, 19 insertions(+), 4 deletions(-) diff --git a/docs/hrx.md b/docs/hrx.md index 9f3bf1e..d5f35ef 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -318,8 +318,15 @@ output identical to the CPU backend. The alignment change is a repro stress knob and is **not** part of the fix. What the fix does not cover: the fault needs router logits that are NaN, and -the fix routes such a lane to a valid expert instead of faulting, so the token -still decodes with the wrong expert. Where the NaNs come from is open +the fix routes such a lane to a valid expert instead of faulting. Since +[llama.cpp #26](https://github.com/1bit-MONSTER/llama.cpp/pull/26) that case is loud: a router row +with no ordered logit publishes NaN, and llama-server refuses to sample NaN logits (it logs +"HRX returned NaN logits" and aborts) instead of decoding a wrong but plausible token. On the +author's repro 6 of 13 samples used to decode `' Paris???…'` silently; they now abort. Healthy +runs are unchanged (Qwen3-Coder-30B and ZAYA1-8B perplexity on HRX0 identical to before). +`GGML_HRX_FA_PARTIAL_ALIGN` (default 4096) sets the decode-split partials' alignment, a test knob +that reproduces the fault's layout; the investigation's artifacts are archived privately. +Where the NaNs come from is open ([engine#140](https://github.com/1bit-MONSTER/engine/issues/140)). One measurement points at the kernel moving the pages behind HRX's buffers: greedy `Qwen3-0.6B` on `HRX0`, 300 identical requests, gave different log-probs on diff --git a/docs/vulkan.md b/docs/vulkan.md index 9c02e82..6e08bc1 100644 --- a/docs/vulkan.md +++ b/docs/vulkan.md @@ -155,6 +155,14 @@ build/hrx/llama/bin/llama-quantize zaya1-8b-f16.gguf zaya1-8b-Q4_K_M.gguf Q4_K_M 1bit serve -m zaya1-8b-Q4_K_M.gguf --device vulkan --ctx-size 8192 ``` +**ZAYA GGUFs converted before 2026-09-27 tokenize chat turns wrongly.** Their `<|im_start|>` (id 105) +and `` were stored as ordinary tokens, so llama.cpp spelled each turn marker out as seven text +tokens and chat answers went off-template ("12 + 30" gave "22"). The converter marks them as +control tokens since [llama.cpp #27](https://github.com/1bit-MONSTER/llama.cpp/pull/27), and every +ZAYA GGUF on the 1bit-MONSTER Hugging Face org was fixed in place the same day (the tensors are +unchanged). Re-download, or reconvert. The teacher-forced checks above never saw it, because they +feed token IDs directly. + Convert with this converter: it writes the grouped convolution's weights tap-major, which the graph expects, so GGUFs made by other converters do not load. diff --git a/registry/architectures.json b/registry/architectures.json index c3ecf20..7fcec4c 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8", - "llama.cpp (hrx)": "895d63f0ed86c1d676a703b01c425717ac484871", + "llama.cpp (hrx)": "fa226f9377a3025a171aea3eb8839c2f9f052d37", "zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb" }, "counts": { diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 895d63f..fa226f9 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 895d63f0ed86c1d676a703b01c425717ac484871 +Subproject commit fa226f9377a3025a171aea3eb8839c2f9f052d37