Skip to content

docs/hrx: root cause of the #123/#140 NaNs (decode-split q8 pack barrier) - #177

Merged
bong-water-water-bong merged 1 commit into
mainfrom
docs/hrx-123-root-cause
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
docs/hrx-123-root-cause

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Docs follow-up to #176. docs/hrx.md still called the NaN source open and pointed at page migration. It now states the root cause (the LDS-only barrier before pack_completed_q8, fixed in llama.cpp 00adc2b via 1bit-MONSTER/llama.cpp#28 and #29) and the measured fix. The decode-split note now says the kernel stays off by default for speed, not correctness: per the existing table it is slower on Qwen3-0.6B and ZAYA1-8B and faster only on Coder-30B. Docs only.

🤖 Generated with Claude Code

…er, fixed in llama.cpp 00adc2b

Replace the open-question paragraph (page migration) with the root cause and the
measured fix, and say decode-split is now off for speed, not correctness.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 27, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 7d3e911

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

176 - Partially compliant

Compliant requirements:

  • third_party/llama.cpp moved from fa226f9 to 00adc2b
  • Root cause fix implemented with global barriers in llama.cpp
  • Measurements provided showing reduction in divergent results and NaN occurrences

Non-compliant requirements:

  • None

Requires further human verification:

  • None

28 - Partially compliant

Compliant requirements:

  • third_party/llama.cpp-vulkan pinned to v0.5.0
  • ONEBIT_VULKAN flag support added
  • Documentation updated with vulkan.md and other relevant docs
  • bump-llama-vulkan.yml workflow added

Non-compliant requirements:

  • None

Requires further human verification:

  • None

29 - Partially compliant

Compliant requirements:

  • third_party/llama.cpp pinned to 1bit/hrx-vulkan-patched branch
  • IQ3_XXS matmul kernel and matcher added
  • Honest op claims implemented to prevent unsupported node dispatching
  • Rebase logic implemented in bump-hrx.yml
  • Previous tip tagged for reachability

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Documentation Clarity

The documentation update correctly identifies and describes the root cause of the NaNs in decode-split flash attention. However, it could be more explicit about how the fix impacts performance trade-offs, especially since the decode-split kernel is slower on some models (like Qwen3-0.6B and ZAYA1-8B) but faster on others (like Qwen3-Coder-30B). Clarifying this in the documentation would help users make informed decisions about enabling or disabling the feature.

- **Decode-split flash attention is off by default.** It used to give nondeterministic attention on
  HRX0 ([#140](https://github.com/1bit-MONSTER/engine/issues/140), up to 3.66 nats apart between
  identical requests on Qwen3-0.6B) and to fault Qwen3-Coder-30B-A3B. That is fixed since llama.cpp
  `00adc2b` (see the #123 section below), but the kernel is only faster on some models, so
  `1bit serve --device hrx` still starts llama-server with `GGML_HRX_DISABLE_DISPATCH=decode_split`.
  The table was measured before the fix:

@bong-water-water-bong
bong-water-water-bong merged commit 97b78aa into main Sep 27, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the docs/hrx-123-root-cause branch September 27, 2026 18:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant