serve: decode-split flash attention back on for --device hrx - #178
Conversation
…verted) #148 turned HRX0's decode-split flash attention off because it gave nondeterministic attention (#140) and faulted MoE models (#123). The cause was the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after the fix (llama-bench, interleaved runs), decode-split is as fast or faster: Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs 47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0. ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the new table. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
…and PORTING status (#182) The decode-split q8 pack read global output behind an LDS-only barrier; fixed in llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised (#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Reverts the #148 default now that #176 fixed the cause (#123/#140).
Post-fix decode tok/s (llama-bench, interleaved runs on strixhalo):
The old docs table (split slower on 0.6B) was measured while the race was live.
Checks: serve_e2e (vulkan, hrx, vulkan+hrx prefill) and registry tests pass. With the default, the launched llama-server has no GGML_HRX_DISABLE_DISPATCH; with
ONEBIT_HRX_DECODE_SPLIT=0it getsdecode_split. Correctness with split on, under forced queue evictions on the #176 build: 0.6B 0/59 divergent, Coder-30B 8/8 with no NaN.Behaviour change:
ONEBIT_HRX_DECODE_SPLIT=0is now the opt-out (it used to be=1to opt in).🤖 Generated with Claude Code