From e6a9eb83ce801bd1f52b52da531385916ae15ef8 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sun, 27 Sep 2026 16:51:06 -0300 Subject: [PATCH] serve: decode-split flash attention back on for --device hrx (#148 reverted) #148 turned HRX0's decode-split flash attention off because it gave nondeterministic attention (#140) and faulted MoE models (#123). The cause was the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after the fix (llama-bench, interleaved runs), decode-split is as fast or faster: Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs 47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0. ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the new table. Co-Authored-By: Claude Opus 5.5 --- app/serve.cpp | 9 ++++----- docs/hrx.md | 33 ++++++++++++++++----------------- 2 files changed, 20 insertions(+), 22 deletions(-) diff --git a/app/serve.cpp b/app/serve.cpp index ee62a72..189a478 100644 --- a/app/serve.cpp +++ b/app/serve.cpp @@ -661,12 +661,11 @@ Launch launch_for(const Options& o, const std::string& device, int child_port) { const std::string hsa = hrx_libhsa(o.hrx_libhsa); if (!hsa.empty()) env.push_back("IREE_HAL_AMDGPU_LIBHSA_PATH=" + hsa); } - // HRX0 decode through flash_attention_decode_split gives nondeterministic, sometimes wrong - // attention and intermittent GPU faults (MoE models); the fallback is deterministic and, on - // most models, faster (#140). A GGML_HRX_DISABLE_DISPATCH the user sets wins, and - // ONEBIT_HRX_DECODE_SPLIT=1 turns the kernel back on, for testing a fix. + // HRX0 decodes through flash_attention_decode_split since its q8 pack race was fixed + // (llama.cpp 00adc2b, #123/#140); it is faster once there is context. A + // GGML_HRX_DISABLE_DISPATCH the user sets wins, and ONEBIT_HRX_DECODE_SPLIT=0 turns it off. const char* keep_split = std::getenv("ONEBIT_HRX_DECODE_SPLIT"); - if (device == "hrx" && !(keep_split && std::string(keep_split) == "1")) + if (device == "hrx" && keep_split && std::string(keep_split) == "0") env.push_back("GGML_HRX_DISABLE_DISPATCH=decode_split"); } else if (device == "ds4") { // DwarfStar: its own GGUF layouts only; it opens its port after the model has loaded diff --git a/docs/hrx.md b/docs/hrx.md index d4becd0..29dd80b 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -122,23 +122,22 @@ Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan): - **Several sequences per batch fail.** `llama-perplexity` with `n_seq` > 1 stops on an unsupported 3-D MUL_MAT; use `-b 512`. - **`-fa off` fails.** A SET_ROWS into the non-flash-attention V cache is rejected. -- **Decode-split flash attention is off by default.** It used to give nondeterministic attention on - HRX0 ([#140](https://github.com/1bit-MONSTER/engine/issues/140), up to 3.66 nats apart between - identical requests on Qwen3-0.6B) and to fault Qwen3-Coder-30B-A3B. That is fixed since llama.cpp - `00adc2b` (see the #123 section below), but the kernel is only faster on some models, so - `1bit serve --device hrx` still starts llama-server with `GGML_HRX_DISABLE_DISPATCH=decode_split`. - The table was measured before the fix: - - | model (Q4_K_M, tg) | with decode-split | without (the default) | - |---|---|---| - | Qwen3-0.6B | 169 ± 34 tok/s, output varies | 322 tok/s, deterministic | - | Qwen3-Coder-30B-A3B | faults 2 of 3 runs (80–88 tok/s when it survives) | 66–71 tok/s, 5 of 5 | - | ZAYA1-8B | 23.5 tok/s | 25.5 tok/s | - - (ZAYA1-8B decodes at 47.9 tok/s since the HRX kernels below.) - - A `GGML_HRX_DISABLE_DISPATCH` you set yourself wins, and `ONEBIT_HRX_DECODE_SPLIT=1` turns the - kernel back on, for testing a fix. +- **Decode-split flash attention is on.** `flash_attention_decode_split_next_q8` used to give + nondeterministic attention on HRX0 ([#140](https://github.com/1bit-MONSTER/engine/issues/140)) and + to fault Qwen3-Coder-30B-A3B, so `1bit serve --device hrx` turned it off (#148). Since llama.cpp + `00adc2b` (see the #123 section below) it is correct, and it is as fast or faster: + + | model (Q4_K_M), tg tok/s | ctx 0 | ctx 512 | ctx 2100 | + |---|---|---|---| + | Qwen3-0.6B, decode-split on | 286 | 254 | 140–158 | + | Qwen3-0.6B, off | 296–306 | 220–223 | 118–122 | + | Qwen3-Coder-30B-A3B, on | 78–85 | 81 | 42–62 | + | Qwen3-Coder-30B-A3B, off | 78–81 | 65–69 | 47–49 | + | ZAYA1-8B, on | 39–41 | 40–43 | 38–42 | + | ZAYA1-8B, off | 39–42 | 39–40 | 34–36 | + + (llama-bench `-p 0 -n 32/64`, runs interleaved on a shared strixhalo, 2026-09-27.) + `ONEBIT_HRX_DECODE_SPLIT=0` turns it off; a `GGML_HRX_DISABLE_DISPATCH` you set yourself wins. The prefill split below is not affected by either: HRX0 only prefills whole ubatches there, and decoding is Vulkan's.