Skip to content

hrx: make the all-NaN router-logit case loud instead of a silent wrong expert (engine#123) - #26

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-moe-router-nan-loud
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-moe-router-nan-loud

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

The branch author's commit message has the full reasoning. In short, after #25 an all-NaN router row no longer faults, but it still decoded silently with a valid-but-wrong expert: the argmax seed is a finite -FLT_MAX, so the softmax gave a plausible uniform weight. This PR makes that case loud.

  • router_top8_f32.loom: a row with no ordered candidate publishes its own NaN, so the corruption reaches the logits.
  • common/sampling.cpp: refuses to sample NaN logits, logging and aborting. NaN is never a legitimate logit, unlike -inf.
  • dispatch-flash-attention.cpp: GGML_HRX_FA_PARTIAL_ALIGN, a test knob for the decode-split partials' alignment. The default stays 4096, so production is unchanged.

Author's repro: Qwen3-Coder-30B, 2,113 context, under memory pressure. Before, 6/13 samples silently decoded ' Paris???????????????'; after, they abort with "HRX returned NaN logits … refusing to sample a silently wrong token".

Review checks (strixhalo, this branch, GGML_HRX_DISABLE_DISPATCH=decode_split as 1bit serve sets, under mem-guard):

  • Perplexity, 8 chunks, is identical to before, so the kernel change only acts on NaN rows: Qwen3-Coder-30B HRX0 17.2526, ZAYA1-8B HRX0 25.2474.
  • llama-server chat, 3 requests each: Qwen3-Coder-30B and ZAYA1-8B on HRX0 and Vulkan0 all answer "Paris", with no false NaN abort.

Trade-off to know about: the NaN check in common/sampling.cpp is shared by every backend and aborts the whole server, dropping all in-flight requests. Fail-stop is the right call for a silent-wrong-token bug, and the check costs one pass over the logits per sampled token. A per-request error would be gentler if this ever fires in production.

🤖 Generated with Claude Code

…g expert

engine#123's residual, after #25 stopped the 0x7FFFFFFF sentinel from faulting: when
the driver migrates a page behind in-flight HRX work (the engine#140-class
nondeterminism), a row's router logits go all-NaN. The argmax seed is a finite
-FLT_MAX, so the row's softmax still produced a plausible uniform 1/route_count and
the token decoded with a valid-but-wrong expert, silently.

- router_top8_f32.loom: when no lane of the row found an ordered candidate, publish
  the row's own NaN instead of that masked uniform weight, so the corruption reaches
  the logits rather than being hidden behind a plausible one.
- common/sampling.cpp: refuse to sample NaN logits - log and abort loudly. NaN is
  never a legitimate logit, unlike -inf, which masking legitimately uses.
- dispatch-flash-attention.cpp: GGML_HRX_FA_PARTIAL_ALIGN selects the decode-split
  partial transients' alignment (default 4096, production unchanged); 256 reproduces
  the engine#123/ggml-org#140 rig oracle without editing stress literals.

Verified on the MoE repro (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, 2113 ctx, fresh
server per sample, under memory pressure): before, 6/13 samples silently decoded
' Paris???????????????'; after, the corrupted samples abort with "HRX returned NaN
logits ... refusing to sample a silently wrong token (engine#123)".
@github-actions github-actions Bot added the ggml label Sep 27, 2026
@bong-water-water-bong
bong-water-water-bong merged commit fdd8f1c into 1bit/hrx-vulkan-patched Sep 27, 2026
10 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant