Skip to content

hrx: never publish the MoE router no-winner sentinel as an expert id - #25

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-moe-router-expert-id
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-moe-router-expert-id

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Summary

router_top8_f32.loom seeds each lane's argmax with best_id = 0x7FFFFFFF and only replaces it when a candidate is strictly greater (ordered ogt) or wins the equal-value tie-break (oeq). When a lane's candidate logits are unordered (NaN) or below -FLT_MAX, neither fires and the sentinel 0x7FFFFFFF is published verbatim into route_ids.

Consumers treat the id as bounded via index.assume [range(...,0,127)] — an optimizer hint, not a runtime check — and compute expert * weight_expert_bytes in 32 bits. 0x7FFFFFFF * 884736 mod 2^32 = 0xFFF28000 (~4.29 GB), so the kernel issues a global read far past the end of the weights buffer:

Warning: Queue error - HSA_STATUS_ERROR_MEMORY_FAULT
llama_decode: failed to decode, ret = -3

Fix

Seed with the lane's own first expert id (lane_expert_base, always < expert_count for a real lane) instead of the sentinel.

This is bit-identical whenever a lane does find a winner: the seed only participates in the first comparison against best_value = -FLT_MAX, and for an equal-valued first candidate the old 0x7FFFFFFF seed and lane_expert_base both resolve to the same id. A no-winner lane now publishes a valid id whose selected_logit stays -FLT_MAX, so its route weight normalizes to 0 and contributes nothing.

Evidence (gfx1151 / Strix Halo, Qwen3-Coder-30B-A3B-Instruct-Q4_K_M, -dev HRX0)

Fault is layout-sensitive, so it is measured with the engine#123 repro rig's 256-byte partial-transient alignment (a stress knob that makes the fault show up; it is not part of this change).

build llama-bench -p 0 -n 8 faults
pre-fix, -d 2100 256 B partials 5/5
this commit, -d 2100 256 B partials 0/5
this commit, -d 3000 / -d 4800 256 B partials 0/3 / 0/3
this commit, -d 1500 (≤2048 path) 256 B partials 0/3

Root cause was pinned with a serial-execution trace (GGML_HRX_DEBUG_SERIAL_EXECUTION=1): the fault occurs in qwen3_moe_routed_gate_up_swiglu_q4k_q8 (command 228), binding blk.37.ffn_gate_exps/blk.37.ffn_up_exps. The logged fault page matched up_binding + 0xFFF28000 + 25*1152 exactly, and giving the weights buffer 8 GB of tail slack (GGML_HRX_ALLOC_PAD_MB=8192) masks it entirely (0/3).

Correctness: greedy llama-server (temperature 0, fixed prompt) on HRX0 produces the same output as the CPU backend of the same build.

Fixes 1bit-MONSTER/engine#123. Note the fault is not in the decode-split multipass path — disabling that dispatch only changed the layout enough to hide it.

router_top8_f32 seeds each lane's argmax with best_id = 0x7FFFFFFF and only
replaces it when a candidate is strictly greater (ordered ogt) or wins the
equal-value tie-break (oeq). When a lane's candidate logits are unordered
(NaN) or below -FLT_MAX, neither fires and the sentinel is published verbatim
into route_ids. Consumers (e.g. qwen3_moe_routed_gate_up_swiglu_q4k_q8) treat
the id as bounded via index.assume -- an optimizer hint, not a check -- and
compute expert * weight_expert_bytes in 32 bits, so 0x7FFFFFFF wraps to
0xFFF28000 (~4.29 GB) and the kernel issues a global read past the end of the
weights buffer: HSA_STATUS_ERROR_MEMORY_FAULT / llama_decode ret=-3.

Seed with the lane's own first expert id instead. lane_expert_base is always
< expert_count for a real lane, the result is bit-identical whenever a lane
does find a winner (the seed only participates in the first comparison), and a
no-winner lane now publishes a valid id whose selected logit is -FLT_MAX, so
its route weight normalizes to 0 and contributes nothing.

Verified on gfx1151 (Strix Halo), Qwen3-Coder-30B-A3B-Instruct-Q4_K_M, HRX0,
llama-bench -p 0 -n 8: -d 2100 0/5, -d 3000 0/3, -d 4800 0/3, <=2048 path
0/3, and greedy output identical to the CPU backend of the same build.

Fixes 1bit-MONSTER/engine#123 (the fault is not in the decode-split multipass
path; disabling that dispatch only changed the layout enough to mask it).
@github-actions github-actions Bot added the ggml label Sep 27, 2026
@bong-water-water-bong
bong-water-water-bong merged commit 895d63f into 1bit/hrx-vulkan-patched Sep 27, 2026
1 check passed
bong-water-water-bong pushed a commit to 1bit-MONSTER/engine that referenced this pull request Sep 27, 2026
… leaves open

Re-points third_party/llama.cpp from the PR head (53c0715) to the merge
commit of 1bit-MONSTER/llama.cpp#25 and regenerates registry/architectures.json
(the registry_pins check failed on the moved pin). docs/hrx.md: the merge
commit, and that the fix turns NaN router logits into a valid-but-wrong expert
instead of a fault; the NaN source stays open (#140), with the page-migration
measurement that points at it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit to 1bit-MONSTER/engine that referenced this pull request Sep 27, 2026
* hrx: bump llama.cpp for the MoE router expert-id fault fix (#123)

Pins third_party/llama.cpp at the fix for engine#123: router_top8_f32 could
publish its 0x7FFFFFFF no-winner seed as an expert id, which
qwen3_moe_routed_gate_up_swiglu_q4k_q8 then turned into a wild weight address
(HSA_STATUS_ERROR_MEMORY_FAULT / llama_decode ret=-3). Upstream PR:
1bit-MONSTER/llama.cpp#25

Re-point at the merge commit once #25 lands.

* hrx: pin llama.cpp at the #25 merge (895d63f); registry; what the fix leaves open

Re-points third_party/llama.cpp from the PR head (53c0715) to the merge
commit of 1bit-MONSTER/llama.cpp#25 and regenerates registry/architectures.json
(the registry_pins check failed on the moved pin). docs/hrx.md: the merge
commit, and that the fix turns NaN router logits into a valid-but-wrong expert
instead of a fault; the NaN source stays open (#140), with the page-migration
measurement that points at it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* registry: regenerate for llama.cpp 895d63f (the #25 merge commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant