Skip to content

serve: decode-split flash attention back on for --device hrx - #178

Merged
bong-water-water-bong merged 1 commit into
mainfrom
hrx-decode-split-default
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
hrx-decode-split-default

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Reverts the #148 default now that #176 fixed the cause (#123/#140).

Post-fix decode tok/s (llama-bench, interleaved runs on strixhalo):

model ctx 0 on / off ctx 512 on / off ctx 2100 on / off
Qwen3-0.6B 286 / 296–306 254 / 220–223 140–158 / 118–122
Qwen3-Coder-30B-A3B 78–85 / 78–81 81 / 65–69 42–62 / 47–49
ZAYA1-8B 39–41 / 39–42 40–43 / 39–40 38–42 / 34–36

The old docs table (split slower on 0.6B) was measured while the race was live.

Checks: serve_e2e (vulkan, hrx, vulkan+hrx prefill) and registry tests pass. With the default, the launched llama-server has no GGML_HRX_DISABLE_DISPATCH; with ONEBIT_HRX_DECODE_SPLIT=0 it gets decode_split. Correctness with split on, under forced queue evictions on the #176 build: 0.6B 0/59 divergent, Coder-30B 8/8 with no NaN.

Behaviour change: ONEBIT_HRX_DECODE_SPLIT=0 is now the opt-out (it used to be =1 to opt in).

🤖 Generated with Claude Code

…verted)

#148 turned HRX0's decode-split flash attention off because it gave
nondeterministic attention (#140) and faulted MoE models (#123). The cause was
the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after
the fix (llama-bench, interleaved runs), decode-split is as fast or faster:
Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs
47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0.

ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and
a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the
new table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 27, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit e6a9eb8

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

148 - Partially compliant

Compliant requirements:

  • 1bit serve --device hrx now starts llama-server without GGML_HRX_DISABLE_DISPATCH=decode_split by default
  • User-set values for GGML_HRX_DISABLE_DISPATCH still win
  • ONEBIT_HRX_DECODE_SPLIT=0 turns the kernel off as intended
  • --prefill-device hrx remains unchanged

Non-compliant requirements:

  • None

Requires further human verification:

  • None

176 - Partially compliant

Compliant requirements:

Non-compliant requirements:

  • None

Requires further human verification:

  • None

123 - Partially compliant

Compliant requirements:

  • The fix for the intermittent faults is implemented in the llama.cpp update
  • The fix involves changing kernel.barrier<workgroup> to kernel.barrier<global> at 5 sites
  • The fix also adds missing global barrier to undispatched standalone reduce_f32 kernels

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Inverted Logic for Decode-Split Control

The logic for controlling the decode_split flash attention kernel has been inverted. Previously, ONEBIT_HRX_DECODE_SPLIT=1 was used to enable the kernel, but now ONEBIT_HRX_DECODE_SPLIT=0 is used to disable it. This change in behavior should be clearly documented and tested to ensure backward compatibility and correct operation.

if (device == "hrx" && keep_split && std::string(keep_split) == "0")
    env.push_back("GGML_HRX_DISABLE_DISPATCH=decode_split");
Documentation Update Required

The documentation in docs/hrx.md needs to be updated to reflect the new default behavior where decode-split flash attention is enabled by default. The previous documentation indicated that it was disabled by default, which is no longer accurate.

- **Decode-split flash attention is on.** `flash_attention_decode_split_next_q8` used to give
  nondeterministic attention on HRX0 ([#140](https://github.com/1bit-MONSTER/engine/issues/140)) and
  to fault Qwen3-Coder-30B-A3B, so `1bit serve --device hrx` turned it off (#148). Since llama.cpp
  `00adc2b` (see the #123 section below) it is correct, and it is as fast or faster:

  | model (Q4_K_M), tg tok/s | ctx 0 | ctx 512 | ctx 2100 |
  |---|---|---|---|
  | Qwen3-0.6B, decode-split on | 286 | 254 | 140–158 |
  | Qwen3-0.6B, off | 296–306 | 220–223 | 118–122 |
  | Qwen3-Coder-30B-A3B, on | 78–85 | 81 | 42–62 |
  | Qwen3-Coder-30B-A3B, off | 78–81 | 65–69 | 47–49 |
  | ZAYA1-8B, on | 39–41 | 40–43 | 38–42 |
  | ZAYA1-8B, off | 39–42 | 39–40 | 34–36 |

  (llama-bench `-p 0 -n 32/64`, runs interleaved on a shared strixhalo, 2026-09-27.)
  `ONEBIT_HRX_DECODE_SPLIT=0` turns it off; a `GGML_HRX_DISABLE_DISPATCH` you set yourself wins.

@bong-water-water-bong
bong-water-water-bong merged commit 615c7ff into main Sep 27, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-decode-split-default branch September 27, 2026 19:53
bong-water-water-bong added a commit that referenced this pull request Sep 28, 2026
…and PORTING status (#182)

The decode-split q8 pack read global output behind an LDS-only barrier; fixed in
llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised
(#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the
NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant