From 71048a3aff0740b5f690193edcf1751f9079cf24 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Sun, 27 Sep 2026 04:09:02 -0300 Subject: [PATCH 1/3] hrx: bump llama.cpp for the MoE router expert-id fault fix (#123) Pins third_party/llama.cpp at the fix for engine#123: router_top8_f32 could publish its 0x7FFFFFFF no-winner seed as an expert id, which qwen3_moe_routed_gate_up_swiglu_q4k_q8 then turned into a wild weight address (HSA_STATUS_ERROR_MEMORY_FAULT / llama_decode ret=-3). Upstream PR: https://github.com/1bit-MONSTER/llama.cpp/pull/25 Re-point at the merge commit once #25 lands. --- docs/hrx.md | 23 +++++++++++++++++++++++ third_party/llama.cpp | 2 +- 2 files changed, 24 insertions(+), 1 deletion(-) diff --git a/docs/hrx.md b/docs/hrx.md index aaed81b..a26de23 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -293,6 +293,29 @@ llama.cpp with ggml-hrx2 on `HRX20`): HRX prefill is now faster than Vulkan's. +## MoE router expert-id fault fix ([engine#123](https://github.com/1bit-MONSTER/engine/issues/123)) + +`llama_decode` could die with `HSA_STATUS_ERROR_MEMORY_FAULT` (`ret = -3`) on MoE +models (`Qwen3-Coder-30B-A3B`). It was reported against the decode-split +multipass path, but it is not in flash attention: the fault is a wild weight +read in `qwen3_moe_routed_gate_up_swiglu_q4k_q8`. + +`router_top8_f32.loom` seeded each lane's argmax with `best_id = 0x7FFFFFFF` and +only replaced it on an ordered `ogt` or the equal-value tie-break. A lane whose +candidate logits are unordered (NaN) or below `-FLT_MAX` therefore published the +sentinel verbatim into `route_ids`; consumers treat the id as bounded via +`index.assume` (a hint, not a check) and compute `expert * expert_stride` in 32 +bits, so `0x7FFFFFFF * 884736 mod 2^32 = 0xFFF28000` (~4.29 GB) walked off the +end of the weights buffer. The fix seeds the argmax with the lane's own first +expert id. Upstream PR: [1bit-MONSTER/llama.cpp#25](https://github.com/1bit-MONSTER/llama.cpp/pull/25). + +Measured on `HRX0` (`Qwen3-Coder-30B-A3B-Instruct-Q4_K_M`, `llama-bench -p 0 -n 8`, +quiet box) with the fault made observable by aligning the three decode-split +*partial* transients to 256 B instead of 4096 B: pre-fix `-d 2100` 5/5 faults, +post-fix `-d 2100` 0/5, `-d 3000`/`-d 4800` 0/3, `<=2048` path 0/3, and greedy +output identical to the CPU backend. The alignment change is a repro stress knob +and is **not** part of the fix. + ## Not yet - **Fast sub-4-bit and IQ4 kernels on HRX.** With our patches the UD files are diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 358cafc..53c0715 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 358cafc249a0234e5c7ffc46201f20aa6682b452 +Subproject commit 53c071525c18a84bd4a0e129b8face623db72e0e From 13774ca6170c1360a689f307ae0bba0772a0bcae Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sun, 27 Sep 2026 04:16:50 -0300 Subject: [PATCH 2/3] hrx: pin llama.cpp at the #25 merge (895d63f); registry; what the fix leaves open Re-points third_party/llama.cpp from the PR head (53c0715) to the merge commit of 1bit-MONSTER/llama.cpp#25 and regenerates registry/architectures.json (the registry_pins check failed on the moved pin). docs/hrx.md: the merge commit, and that the fix turns NaN router logits into a valid-but-wrong expert instead of a fault; the NaN source stays open (#140), with the page-migration measurement that points at it. Co-Authored-By: Claude Opus 5.5 --- docs/hrx.md | 15 ++++++++++++++- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 3 files changed, 16 insertions(+), 3 deletions(-) diff --git a/docs/hrx.md b/docs/hrx.md index a26de23..9f3bf1e 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -307,7 +307,8 @@ sentinel verbatim into `route_ids`; consumers treat the id as bounded via `index.assume` (a hint, not a check) and compute `expert * expert_stride` in 32 bits, so `0x7FFFFFFF * 884736 mod 2^32 = 0xFFF28000` (~4.29 GB) walked off the end of the weights buffer. The fix seeds the argmax with the lane's own first -expert id. Upstream PR: [1bit-MONSTER/llama.cpp#25](https://github.com/1bit-MONSTER/llama.cpp/pull/25). +expert id. Upstream PR: [1bit-MONSTER/llama.cpp#25](https://github.com/1bit-MONSTER/llama.cpp/pull/25), +merged as `895d63f`, which the engine pins. Measured on `HRX0` (`Qwen3-Coder-30B-A3B-Instruct-Q4_K_M`, `llama-bench -p 0 -n 8`, quiet box) with the fault made observable by aligning the three decode-split @@ -316,6 +317,18 @@ post-fix `-d 2100` 0/5, `-d 3000`/`-d 4800` 0/3, `<=2048` path 0/3, and greedy output identical to the CPU backend. The alignment change is a repro stress knob and is **not** part of the fix. +What the fix does not cover: the fault needs router logits that are NaN, and +the fix routes such a lane to a valid expert instead of faulting, so the token +still decodes with the wrong expert. Where the NaNs come from is open +([engine#140](https://github.com/1bit-MONSTER/engine/issues/140)). One +measurement points at the kernel moving the pages behind HRX's buffers: greedy +`Qwen3-0.6B` on `HRX0`, 300 identical requests, gave different log-probs on +24-38 % of them with the box's defaults (KSM on, proactive compaction 20, +transparent huge pages `always`), 0-0.7 % with KSM and compaction off, 8-11 % +with proactive compaction alone, 4 % with KSM alone, and 0-0.7 % with every +setting at its default but the server `mlockall`ed and +`vm.compact_unevictable_allowed=0` (2026-09-27, the pin before this fix). + ## Not yet - **Fast sub-4-bit and IQ4 kernels on HRX.** With our patches the UD files are diff --git a/registry/architectures.json b/registry/architectures.json index 9ad3d87..a0b4ca3 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8", - "llama.cpp (hrx)": "358cafc249a0234e5c7ffc46201f20aa6682b452", + "llama.cpp (hrx)": "53c071525c18a84bd4a0e129b8face623db72e0e", "zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb" }, "counts": { diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 53c0715..895d63f 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 53c071525c18a84bd4a0e129b8face623db72e0e +Subproject commit 895d63f0ed86c1d676a703b01c425717ac484871 From df519c8ef96bfb77d08fb4687891af4c30524b1c Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sun, 27 Sep 2026 04:29:33 -0300 Subject: [PATCH 3/3] registry: regenerate for llama.cpp 895d63f (the #25 merge commit) Co-Authored-By: Claude Opus 5.5 --- registry/architectures.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/registry/architectures.json b/registry/architectures.json index a0b4ca3..c3ecf20 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8", - "llama.cpp (hrx)": "53c071525c18a84bd4a0e129b8face623db72e0e", + "llama.cpp (hrx)": "895d63f0ed86c1d676a703b01c425717ac484871", "zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb" }, "counts": {