Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 9 additions & 2 deletions docs/hrx.md
Original file line number Diff line number Diff line change
Expand Up @@ -318,8 +318,15 @@ output identical to the CPU backend. The alignment change is a repro stress knob
and is **not** part of the fix.

What the fix does not cover: the fault needs router logits that are NaN, and
the fix routes such a lane to a valid expert instead of faulting, so the token
still decodes with the wrong expert. Where the NaNs come from is open
the fix routes such a lane to a valid expert instead of faulting. Since
[llama.cpp #26](https://github.com/1bit-MONSTER/llama.cpp/pull/26) that case is loud: a router row
with no ordered logit publishes NaN, and llama-server refuses to sample NaN logits (it logs
"HRX returned NaN logits" and aborts) instead of decoding a wrong but plausible token. On the
author's repro 6 of 13 samples used to decode `' Paris???…'` silently; they now abort. Healthy
runs are unchanged (Qwen3-Coder-30B and ZAYA1-8B perplexity on HRX0 identical to before).
`GGML_HRX_FA_PARTIAL_ALIGN` (default 4096) sets the decode-split partials' alignment, a test knob
that reproduces the fault's layout; the investigation's artifacts are archived privately.
Where the NaNs come from is open
([engine#140](https://github.com/1bit-MONSTER/engine/issues/140)). One
measurement points at the kernel moving the pages behind HRX's buffers: greedy
`Qwen3-0.6B` on `HRX0`, 300 identical requests, gave different log-probs on
Expand Down
8 changes: 8 additions & 0 deletions docs/vulkan.md
Original file line number Diff line number Diff line change
Expand Up @@ -155,6 +155,14 @@ build/hrx/llama/bin/llama-quantize zaya1-8b-f16.gguf zaya1-8b-Q4_K_M.gguf Q4_K_M
1bit serve -m zaya1-8b-Q4_K_M.gguf --device vulkan --ctx-size 8192
```

**ZAYA GGUFs converted before 2026-09-27 tokenize chat turns wrongly.** Their `<|im_start|>` (id 105)
and `<eos>` were stored as ordinary tokens, so llama.cpp spelled each turn marker out as seven text
tokens and chat answers went off-template ("12 + 30" gave "22"). The converter marks them as
control tokens since [llama.cpp #27](https://github.com/1bit-MONSTER/llama.cpp/pull/27), and every
ZAYA GGUF on the 1bit-MONSTER Hugging Face org was fixed in place the same day (the tensors are
unchanged). Re-download, or reconvert. The teacher-forced checks above never saw it, because they
feed token IDs directly.

Convert with this converter: it writes the grouped convolution's weights tap-major, which the
graph expects, so GGUFs made by other converters do not load.

Expand Down
2 changes: 1 addition & 1 deletion registry/architectures.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
"llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8",
"llama.cpp (hrx)": "895d63f0ed86c1d676a703b01c425717ac484871",
"llama.cpp (hrx)": "fa226f9377a3025a171aea3eb8839c2f9f052d37",
"zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb"
},
"counts": {
Expand Down
Loading