Skip to content

Pin llama.cpp 00adc2b: HRX decode-split publishes next_q8 after a global barrier (#123, #140) - #176

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin/hrx-decode-split-barrier
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin/hrx-decode-split-barrier

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp fa226f9 → 00adc2b (fast-forward), bringing 1bit-MONSTER/llama.cpp#28 and #29: the root-cause fix for #123 and #140.

Bug: in every HRX decode-split reduce_fused variant (direct, cooperative, multipass; loom-libs/ops and qwen_moe/qwen3_moe corpora), the barrier between the reduce's global output stores and pack_completed_q8's output loads was kernel.barrier<workgroup>, which orders LDS only (s_waitcnt lgkmcnt(0); s_barrier, no vscnt wait, no buffer_gl0_inv). The pack step could quantise stale output into the next layer's Q8 input. The f32 result was right and the Q8 copy was wrong. Preemption (any process's KFD queue eviction), contention and transient alignment only widened the window. The fix is kernel.barrier<global> at those 5 sites, plus the LDS barrier kept next to it (#29: a Loom barrier fences only the memory space it names). #29 also adds the missing global barrier to the undispatched standalone reduce_f32 kernels.

Evidence (llama.cpp#28, logs in strixhalo ~/wt/hrx123-logs):

  • Qwen3-0.6B under another process's queue evictions: 0/354 divergent after the fix (before: 13–34 per 177).
  • Same with mlockall + forced compaction: 0/118.
  • Qwen3-Coder-30B at GGML_HRX_FA_PARTIAL_ALIGN=256 (the HRX: decode-split multipass path intermittently faults (HSA_STATUS_ERROR_MEMORY_FAULT) above capacity 2048 #123 case): 16/16 correct, 0 NaN (before: 1/8 correct, NaN guard fired 7×).
  • Merged head 9ca9603: 0.6B 0/59, 30B 8/8.
  • Final head 00adc2b: 0.6B under evictions 0/118, 30B at 256 alignment 8/8 with 0 NaN; every variant's ISA has vscnt(0), s_barrier and gl1/gl0_inv before the pack load; decode speed within noise.

The #170 router change and #175's loud-NaN guard remain as defence in depth. Checks on this PR's head (3411558, llama.cpp 00adc2b), strixhalo, built with ONEBIT_HRX=ON ONEBIT_VULKAN=ON:

🤖 Generated with Claude Code

@context7

context7 Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 3411558

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 3411558)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

28 - Partially compliant

Compliant requirements:

  • Pin third_party/llama.cpp-vulkan to ggml-org/llama.cpp v0.5.0 (7fe450e19305)
  • Build llama-server with ONEBIT_VULKAN flag
  • 1bit serve --device vulkan prefers the vulkan build
  • --device hrx keeps the HRX build
  • bump-llama-vulkan.yml creates a daily PR when upstream publishes a release
  • Update docs (docs/vulkan.md; serve.md, hrx.md, README and NOTICE)

Non-compliant requirements:

  • None

Requires further human verification:

  • None

29 - Partially compliant

Compliant requirements:

  • Move third_party/llama.cpp from AMD's commit to 1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched (d2a9239f9)
  • Add IQ3_XXS matmul on HRX with kernel kernels/hrx/mul_mat_vec_iq3xxs_f32.loom and matcher common/dispatch-mul-mat-iq3-xxs.{h,cpp}
  • Add honest op claims to prevent unsupported node dispatches
  • bump-hrx.yml rebases our commits onto each new AMD pair, tagging previous tip patched-<sha>
  • Verify on Strix Halo with Qwen3-0.6B and perplexity tests

Non-compliant requirements:

  • None

Requires further human verification:

  • None

123 - Partially compliant

Compliant requirements:

  • Fix intermittent fault in HRX decode-split multipass path (HSA_STATUS_ERROR_MEMORY_FAULT)
  • Ensure that the barrier between reduce's global output stores and pack_completed_q8's output loads is kernel.barrier<global>
  • Fix the issue where stale output could be quantized into the next layer's Q8 input
  • Validate fix with Qwen3-0.6B under queue evictions and other stress tests

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Subproject Commit Update

The subproject commit for third_party/llama.cpp has been updated from fa226f9377a3025a171aea3eb8839c2f9f052d37 to 00adc2b9d6676f39cb3fcca8b5ad97621e653fae. This update brings in the fixes for the HRX decode-split multipass path fault and the vulkan pinning changes. Ensure that this commit is correctly pinned and that all related changes are properly integrated.

Subproject commit 00adc2b9d6676f39cb3fcca8b5ad97621e653fae
Architecture Source Update

The architecture source for llama.cpp (hrx) has been updated in registry/architectures.json to reflect the new pinned commit 00adc2b9d6676f39cb3fcca8b5ad97621e653fae. This ensures that the registry correctly identifies the source of the HRX backend.

"llama.cpp (hrx)": "00adc2b9d6676f39cb3fcca8b5ad97621e653fae",

…bal barrier (#123, #140)

Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier
between the reduce's global output stores and pack_completed_q8's output loads was
kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output.
It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe
corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to
the undispatched standalone reduce_f32 kernels). The pin moves
fa226f9 -> 00adc2b, fast-forward.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong bong-water-water-bong changed the title Pin llama.cpp 9ca9603: HRX decode-split publishes next_q8 after a global barrier (#123, #140) Pin llama.cpp 00adc2b: HRX decode-split publishes next_q8 after a global barrier (#123, #140) Sep 27, 2026
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 3411558

@bong-water-water-bong
bong-water-water-bong merged commit f4b8cc9 into main Sep 27, 2026
9 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin/hrx-decode-split-barrier branch September 27, 2026 18:42
bong-water-water-bong added a commit that referenced this pull request Sep 27, 2026
…verted) (#178)

#148 turned HRX0's decode-split flash attention off because it gave
nondeterministic attention (#140) and faulted MoE models (#123). The cause was
the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after
the fix (llama-bench, interleaved runs), decode-split is as fast or faster:
Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs
47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0.

ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and
a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the
new table.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 28, 2026
…and PORTING status (#182)

The decode-split q8 pack read global output behind an LDS-only barrier; fixed in
llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised
(#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the
NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant