Pin llama.cpp 00adc2b: HRX decode-split publishes next_q8 after a global barrier (#123, #140) - #176
Conversation
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍(Review updated until commit 3411558)Here are some key observations to aid the review process:
|
…bal barrier (#123, #140) Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier between the reduce's global output stores and pack_completed_q8's output loads was kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output. It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to the undispatched standalone reduce_f32 kernels). The pin moves fa226f9 -> 00adc2b, fast-forward. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2582f50 to
3411558
Compare
|
Persistent review updated to latest commit 3411558 |
…verted) (#178) #148 turned HRX0's decode-split flash attention off because it gave nondeterministic attention (#140) and faulted MoE models (#123). The cause was the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after the fix (llama-bench, interleaved runs), decode-split is as fast or faster: Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs 47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0. ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the new table. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…and PORTING status (#182) The decode-split q8 pack read global output behind an LDS-only barrier; fixed in llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised (#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Moves
third_party/llama.cppfa226f9 → 00adc2b (fast-forward), bringing 1bit-MONSTER/llama.cpp#28 and #29: the root-cause fix for #123 and #140.Bug: in every HRX decode-split
reduce_fusedvariant (direct, cooperative, multipass;loom-libs/opsandqwen_moe/qwen3_moecorpora), the barrier between the reduce's globaloutputstores andpack_completed_q8'soutputloads waskernel.barrier<workgroup>, which orders LDS only (s_waitcnt lgkmcnt(0); s_barrier, novscntwait, nobuffer_gl0_inv). The pack step could quantise staleoutputinto the next layer's Q8 input. The f32 result was right and the Q8 copy was wrong. Preemption (any process's KFD queue eviction), contention and transient alignment only widened the window. The fix iskernel.barrier<global>at those 5 sites, plus the LDS barrier kept next to it (#29: a Loom barrier fences only the memory space it names). #29 also adds the missing global barrier to the undispatched standalonereduce_f32kernels.Evidence (llama.cpp#28, logs in strixhalo
~/wt/hrx123-logs):GGML_HRX_FA_PARTIAL_ALIGN=256(the HRX: decode-split multipass path intermittently faults (HSA_STATUS_ERROR_MEMORY_FAULT) above capacity 2048 #123 case): 16/16 correct, 0 NaN (before: 1/8 correct, NaN guard fired 7×).vscnt(0),s_barrierandgl1/gl0_invbefore the pack load; decode speed within noise.The #170 router change and #175's loud-NaN guard remain as defence in depth. Checks on this PR's head (3411558, llama.cpp 00adc2b), strixhalo, built with ONEBIT_HRX=ON ONEBIT_VULKAN=ON:
🤖 Generated with Claude Code