HRX: fix the MoE router expert-id fault (engine#123) - #170
Conversation
Pins third_party/llama.cpp at the fix for engine#123: router_top8_f32 could publish its 0x7FFFFFFF no-winner seed as an expert id, which qwen3_moe_routed_gate_up_swiglu_q4k_q8 then turned into a wild weight address (HSA_STATUS_ERROR_MEMORY_FAULT / llama_decode ret=-3). Upstream PR: 1bit-MONSTER/llama.cpp#25 Re-point at the merge commit once #25 lands.
|
Docs7 for 1bit-monster/engine
Commit |
… leaves open Re-points third_party/llama.cpp from the PR head (53c0715) to the merge commit of 1bit-MONSTER/llama.cpp#25 and regenerates registry/architectures.json (the registry_pins check failed on the moved pin). docs/hrx.md: the merge commit, and that the fix turns NaN router logits into a valid-but-wrong expert instead of a fault; the NaN source stays open (#140), with the page-migration measurement that points at it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Reviewed. The root cause and fix hold up: seeding with
Worth a follow-up: now that the fault is traced to the router rather than flash attention, #148's decode-split default can be revisited. #140's nondeterminism (up to 3.66 nats between identical requests) is a separate issue and still has to be ruled out with repeated runs before the kernel goes back on. |
What
Bumps
third_party/llama.cppto the fix for #123 and documents it indocs/hrx.md.Upstream: 1bit-MONSTER/llama.cpp#25
Root cause
llama_decodeintermittently died withHSA_STATUS_ERROR_MEMORY_FAULT(ret = -3) on MoE models. It was reported against the decode-split multipass path, but it is not in flash attention — disabling that dispatch only changed the layout enough to mask it. A serial-execution trace put the fault inqwen3_moe_routed_gate_up_swiglu_q4k_q8:router_top8_f32.loomseeded each lane's argmax withbest_id = 0x7FFFFFFFand only replaced it on an orderedogtor the equal-value tie-break. A lane whose candidate logits are unordered (NaN) or below-FLT_MAXpublished the sentinel verbatim intoroute_ids; consumers treat the id as bounded viaindex.assume(a hint, not a check) and computeexpert * expert_stridein 32 bits, so0x7FFFFFFF * 884736 mod 2^32 = 0xFFF28000(~4.29 GB) walked off the end of the weights buffer. The logged fault page matchedup_binding + 0xFFF28000 + 25*1152exactly.Verified
gfx1151(Strix Halo),Qwen3-Coder-30B-A3B-Instruct-Q4_K_M,-dev HRX0,llama-bench -p 0 -n 8, quiet box, on the superset tree withGGML_HRX=ON:-d 2100: 5/5 faults → post-fix 0/5-d 3000/-d 4800: 0/3 / 0/3<=2048path (-d 1500): 0/3llama-serveroutput on HRX0 == CPU backend of the same buildRe-point the pin at the PR merge commit once #25 lands.