From 53c071525c18a84bd4a0e129b8face623db72e0e Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Sun, 27 Sep 2026 04:07:23 -0300 Subject: [PATCH] hrx: never publish the MoE router no-winner sentinel as an expert id router_top8_f32 seeds each lane's argmax with best_id = 0x7FFFFFFF and only replaces it when a candidate is strictly greater (ordered ogt) or wins the equal-value tie-break (oeq). When a lane's candidate logits are unordered (NaN) or below -FLT_MAX, neither fires and the sentinel is published verbatim into route_ids. Consumers (e.g. qwen3_moe_routed_gate_up_swiglu_q4k_q8) treat the id as bounded via index.assume -- an optimizer hint, not a check -- and compute expert * weight_expert_bytes in 32 bits, so 0x7FFFFFFF wraps to 0xFFF28000 (~4.29 GB) and the kernel issues a global read past the end of the weights buffer: HSA_STATUS_ERROR_MEMORY_FAULT / llama_decode ret=-3. Seed with the lane's own first expert id instead. lane_expert_base is always < expert_count for a real lane, the result is bit-identical whenever a lane does find a winner (the seed only participates in the first comparison), and a no-winner lane now publishes a valid id whose selected logit is -FLT_MAX, so its route weight normalizes to 0 and contributes nothing. Verified on gfx1151 (Strix Halo), Qwen3-Coder-30B-A3B-Instruct-Q4_K_M, HRX0, llama-bench -p 0 -n 8: -d 2100 0/5, -d 3000 0/3, -d 4800 0/3, <=2048 path 0/3, and greedy output identical to the CPU backend of the same build. Fixes 1bit-MONSTER/engine#123 (the fault is not in the decode-split multipass path; disabling that dispatch only changed the layout enough to mask it). --- .../kernels/qwen_moe/qwen3_moe/router_top8_f32.loom | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/ggml/src/ggml-hrx/kernel-corpus/kernels/qwen_moe/qwen3_moe/router_top8_f32.loom b/ggml/src/ggml-hrx/kernel-corpus/kernels/qwen_moe/qwen3_moe/router_top8_f32.loom index 1ecccc71ba44..7c2f77fae093 100644 --- a/ggml/src/ggml-hrx/kernel-corpus/kernels/qwen_moe/qwen3_moe/router_top8_f32.loom +++ b/ggml/src/ggml-hrx/kernel-corpus/kernels/qwen_moe/qwen3_moe/router_top8_f32.loom @@ -52,8 +52,14 @@ template.def<@qwen3_moe.router.top8.row> device requires [#target.subgroup.size< %route_ids_view = buffer.view %route_ids_noalias[%c0_offset] : buffer -> view<[%route_id_storage_count]xi32> %route_weights_view = buffer.view %route_weights_noalias[%c0_offset] : buffer -> view<[%route_weight_storage_count]xf32> %initial_logits = vector.load %logits_view[%safe_token, %lane_expert_base] : view<[%launch_token_count]x[%expert_count]xf32> -> vector<[%experts_per_lane]xf32> + // Seed the per-lane argmax with this lane's first expert id instead of the + // 0x7FFFFFFF sentinel. If every candidate is unordered (NaN) or below + // -FLT_MAX, `ogt`/`oeq` never fire and the old seed was published verbatim + // into route_ids, where the consumer turned it into a wild weight address + // (engine#123). lane_expert_base is always < expert_count for a real lane. + %lane_expert_base_i32 = index.cast %lane_expert_base : index to i32 %remaining_final, %selected_logit = scf.for %route = [%c0 to %route_count step %c1](%remaining_logits = %initial_logits : vector<[%experts_per_lane]xf32>, %lane_selected_logit = %negative_large : f32) -> (vector<[%experts_per_lane]xf32>, f32) { - %local_value, %local_id = scf.for %slot = [%c0 to %experts_per_lane step %c1](%best_value = %negative_large : f32, %best_id = %largest_i32 : i32) -> (f32, i32) unroll { + %local_value, %local_id = scf.for %slot = [%c0 to %experts_per_lane step %c1](%best_value = %negative_large : f32, %best_id = %lane_expert_base_i32 : i32) -> (f32, i32) unroll { %candidate_value = vector.extract %remaining_logits[%slot] : vector<[%experts_per_lane]xf32> -> f32 %candidate_expert = index.add %lane_expert_base, %slot : index %candidate_id = index.cast %candidate_expert : index to i32