hrx: make the all-NaN router-logit case loud instead of a silent wrong expert (engine#123) - #26
Merged
bong-water-water-bong merged 1 commit intoSep 27, 2026
Conversation
…g expert engine#123's residual, after #25 stopped the 0x7FFFFFFF sentinel from faulting: when the driver migrates a page behind in-flight HRX work (the engine#140-class nondeterminism), a row's router logits go all-NaN. The argmax seed is a finite -FLT_MAX, so the row's softmax still produced a plausible uniform 1/route_count and the token decoded with a valid-but-wrong expert, silently. - router_top8_f32.loom: when no lane of the row found an ordered candidate, publish the row's own NaN instead of that masked uniform weight, so the corruption reaches the logits rather than being hidden behind a plausible one. - common/sampling.cpp: refuse to sample NaN logits - log and abort loudly. NaN is never a legitimate logit, unlike -inf, which masking legitimately uses. - dispatch-flash-attention.cpp: GGML_HRX_FA_PARTIAL_ALIGN selects the decode-split partial transients' alignment (default 4096, production unchanged); 256 reproduces the engine#123/ggml-org#140 rig oracle without editing stress literals. Verified on the MoE repro (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, 2113 ctx, fresh server per sample, under memory pressure): before, 6/13 samples silently decoded ' Paris???????????????'; after, the corrupted samples abort with "HRX returned NaN logits ... refusing to sample a silently wrong token (engine#123)".
bong-water-water-bong
merged commit Sep 27, 2026
fdd8f1c
into
1bit/hrx-vulkan-patched
10 of 24 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The branch author's commit message has the full reasoning. In short, after #25 an all-NaN router row no longer faults, but it still decoded silently with a valid-but-wrong expert: the argmax seed is a finite
-FLT_MAX, so the softmax gave a plausible uniform weight. This PR makes that case loud.router_top8_f32.loom: a row with no ordered candidate publishes its own NaN, so the corruption reaches the logits.common/sampling.cpp: refuses to sample NaN logits, logging and aborting. NaN is never a legitimate logit, unlike-inf.dispatch-flash-attention.cpp:GGML_HRX_FA_PARTIAL_ALIGN, a test knob for the decode-split partials' alignment. The default stays 4096, so production is unchanged.Author's repro: Qwen3-Coder-30B, 2,113 context, under memory pressure. Before, 6/13 samples silently decoded
' Paris???????????????'; after, they abort with "HRX returned NaN logits … refusing to sample a silently wrong token".Review checks (strixhalo, this branch,
GGML_HRX_DISABLE_DISPATCH=decode_splitas1bit servesets, under mem-guard):llama-serverchat, 3 requests each: Qwen3-Coder-30B and ZAYA1-8B on HRX0 and Vulkan0 all answer "Paris", with no false NaN abort.Trade-off to know about: the NaN check in
common/sampling.cppis shared by every backend and aborts the whole server, dropping all in-flight requests. Fail-stop is the right call for a silent-wrong-token bug, and the check costs one pass over the logits per sampled token. A per-request error would be gentler if this ever fires in production.🤖 Generated with Claude Code