From 54fe30e6aa6cd2bcf953afef0da9216268601348 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong Date: Sun, 27 Sep 2026 23:43:04 -0300 Subject: [PATCH] Blog: milestone - the HRX decode race is solved (#123, #140); README and PORTING status The decode-split q8 pack read global output behind an LDS-only barrier; fixed in llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised (#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md. Co-Authored-By: Claude Opus 5.5 --- README.md | 2 +- blog/2026-09-27-hrx-decode-race-solved.md | 80 +++++++++++++++++++++++ docs/PORTING.md | 2 +- 3 files changed, 82 insertions(+), 2 deletions(-) create mode 100644 blog/2026-09-27-hrx-decode-race-solved.md diff --git a/README.md b/README.md index 16b44c5..73c7a96 100644 --- a/README.md +++ b/README.md @@ -45,7 +45,7 @@ Windows and 1bit OS ([docs/releases.md](docs/releases.md)). > [1bit-MONSTER/lemonade](https://github.com/1bit-MONSTER/lemonade)) passes Lemonade's own LLM test > suite on Vulkan and HRX. Following geramyL's review the engine no longer > vendors Lemonade: Lemonade is the host, 1bit is the engine inside it. Ported so far: HRX on AMD's live ggml-hrx -> ([docs/hrx.md](docs/hrx.md)), Vulkan from upstream llama.cpp's latest release ([docs/vulkan.md](docs/vulkan.md)), the NPU engine on full ELFs with the +> ([docs/hrx.md](docs/hrx.md); its decode-split race, #123/#140, is fixed and the kernel is on by default), Vulkan from upstream llama.cpp's latest release ([docs/vulkan.md](docs/vulkan.md)), the NPU engine on full ELFs with the > upstream XDNA stack pinned ([docs/npu.md](docs/npu.md); its layer kernel is not yet built from > source), ZINC ([docs/zinc.md](docs/zinc.md)) and MLX ([docs/apple.md](docs/apple.md)). ZAYA1-8B (Zyphra) > runs from our llama.cpp on Vulkan, HRX and ROCm, matching transformers, at 93 tok/s decode in diff --git a/blog/2026-09-27-hrx-decode-race-solved.md b/blog/2026-09-27-hrx-decode-race-solved.md new file mode 100644 index 0000000..161fdeb --- /dev/null +++ b/blog/2026-09-27-hrx-decode-race-solved.md @@ -0,0 +1,80 @@ + +tags: hrx, milestone, kernels + +# Milestone: the HRX decode race is solved + +Yesterday's post ended with a workaround. HRX's split flash-attention decode kernel gave +different answers to identical requests, and on Qwen3-Coder-30B it sometimes produced NaNs that +faulted the GPU ([#123](https://github.com/1bit-MONSTER/engine/issues/123), +[#140](https://github.com/1bit-MONSTER/engine/issues/140)). So `1bit serve --device hrx` stopped +using it. Today we found the cause and fixed it, and the kernel is back on by default. It is +correct again, and faster than the path we used without it. + +## What it was + +The decode-split kernel attends over the KV cache in blocks, and a final workgroup combines the +blocks. That workgroup writes the attention output to global memory. Other waves then read it +back to quantize it (`next_q8`) for the next projection. Between the write and the read sat a +workgroup barrier that fences shared memory (LDS) only. It does not wait for global stores or +invalidate the cache the readers use. So a reading wave could pick up the output slot's +*previous* contents. The f32 output was right, the quantized copy was wrong, and when the stale +bytes happened to decode as NaN the MoE router saw NaN logits. + +The window only opens when the GPU preempts the dispatch in the middle. Any process's KFD queue +eviction does that: page compaction, KSM, memory pressure. That is why the failure rate followed +those system settings. For a while it looked like a page-migration bug, but migration was only +the trigger. The fix fences both memory spaces at all five places the kernel does this: +`barrier`, then `barrier` +([docs/hrx.md](../docs/hrx.md#moe-router-expert-id-fault-fix-engine123)). + +## How we know + +A separate process forced queue evictions while the engine served, to make the rare case common: + +- **Qwen3-Coder-30B-A3B** (2113-token prompt): 1 of 8 correct before, 8 of 8 after, with no NaN. +- **Qwen3-0.6B**, greedy: 13-34 divergent answers per ~177 requests before, none in 354 after. + +The last clue came from hashing every kernel's output in serial execution. In every bad request, +the first command whose output diverged was this kernel, with the f32 output identical and the +quantized output different. + +## Faster with it on + +Decode tok/s (Q4_K_M, llama-bench, interleaved on a shared box): + +| model | ctx 2100, split on | ctx 2100, split off | +|---|---|---| +| Qwen3-0.6B | 140-158 | 118-122 | +| Qwen3-Coder-30B-A3B | 42-62 | 47-49 | +| ZAYA1-8B | 38-42 | 34-36 | + +The kernel's multi-pass output step was also vectorised +([#180](https://github.com/1bit-MONSTER/engine/pull/180)): +30-37% at depth 2100, with +bit-identical perplexity. `ONEBIT_HRX_DECODE_SPLIT=0` turns the kernel off if you need to compare. + +## Also this week + +- **Architecture gaps closed.** OPT, GPT-Neo, CodeGen and GPT-J, which upstream llama.cpp has no + model code for, and Zyphra's whole family, from ZAYA1-74B to the vision models, all run from our + llama.cpp. The registry now maps 323 Hugging Face architectures, up from 265 at step 5 + ([docs/registry.md](../docs/registry.md)). +- **GGUF on the NPU**: a repacked GGUF serves on the NPU. Qwen2.5-7B, MiniCPM4-8B and + MiniCPM5-1B answer there + ([#179](https://github.com/1bit-MONSTER/engine/pull/179), [docs/npu.md](../docs/npu.md#from-a-gguf)). + +Packages with the fix ship on Sunday with the weekly release. diff --git a/docs/PORTING.md b/docs/PORTING.md index f5b499b..b227944 100644 --- a/docs/PORTING.md +++ b/docs/PORTING.md @@ -22,7 +22,7 @@ Each step below is one PR (or a short series) that builds and runs on Strix Halo | Step | Component | Source in 1bit-MONSTER | Done when | |---|---|---|---| | 1 | **The engine runs inside Lemonade.** Lemonade runs `1bit serve` as a backend: one model per process behind an OpenAI-compatible API | `tools/unified_server.cpp` | **landed** ([docs/serve.md](serve.md)): NPU, Vulkan, HRX and ZINC pass `serve_e2e` on Strix Halo; `smoke_serve` runs in CI. The engine first embedded Lemonade (v11.9.0 plus local recipes); after geramyL's review it is embedded *into* Lemonade instead: the vendored copy is removed, and the Lemonade recipe that runs `1bit serve` (`onebit`) lives in our fork [1bit-MONSTER/lemonade](https://github.com/1bit-MONSTER/lemonade) ([docs/lemonade.md](lemonade.md)) | -| 2 | **HRX with Vulkan.** llama.cpp with `GGML_HRX=ON` and `GGML_VULKAN=ON` in one build, on AMD's tested pair | `1bit-MONSTER/llama.cpp` `1bit/hrx-vulkan-patched` (AMD's `hrx-graph-develop-v2` plus our commits) + `ROCm/hrx-system`, pinned from `ROCm/ggml-staging-automation` | **landed** ([docs/hrx.md](hrx.md)): `1bit serve` runs the same checkpoint on `Vulkan0` and `HRX0`; kept current by `bump-hrx.yml` | +| 2 | **HRX with Vulkan.** llama.cpp with `GGML_HRX=ON` and `GGML_VULKAN=ON` in one build, on AMD's tested pair | `1bit-MONSTER/llama.cpp` `1bit/hrx-vulkan-patched` (AMD's `hrx-graph-develop-v2` plus our commits) + `ROCm/hrx-system`, pinned from `ROCm/ggml-staging-automation` | **landed** ([docs/hrx.md](hrx.md)): `1bit serve` runs the same checkpoint on `Vulkan0` and `HRX0`; kept current by `bump-hrx.yml`. The decode-split flash-attention race (#123, #140) is fixed (llama.cpp `00adc2b`, #176) and the kernel is on by default again (#178) | | 3 | **NPU engine.** Full ELFs only, and the open 16-tile layer kernel built from source | `engine/npu` (`npu_engine_universal.cpp`, `I8Ctx::init_elf`); ELF dispatch table on `backup/iso-build-elf-native-2026-09-22`; kernel on `bench/fastlane-16tile-corrections-2026-09-22` | **3a–3c landed** ([docs/npu.md](npu.md)). 3a (#8): full ELFs generated in C++, reproducing all 5,153 captured contexts. 3b/3c (#9): the lane runtime (logits bit-identical to the reference lane, 24/24 steps, 11.0 ms/token), tokenizer, `1bit unified`, and Qwen3-0.6B served on the NPU (now through `1bit serve`). The XDNA driver and XRT are pinned upstream and built privately (#12). **Open:** the layer kernel and lm-head artifacts are not yet built from source (npu.md, "Open") | | 3x | **35B MoE on the NPU (experimental, closed source).** Qwen3.6-35B-A3B on the NPU | the private `1bit-MONSTER/npu-kernels` repository | **landed as a private add-on** ([docs/npu.md](npu.md#private-routes)): built into `1bit` with `-DONEBIT_NPU_PRIVATE`; parity against the fp64 reference passes (3 positions, argmax 846 / 198 / 3710), 16.3-16.5 tok/s decode, and `1bit serve --device npu` answers "The capital of France is Paris." Without the add-on, `serve` says the route is not part of the build | | + | **Linux kernel.** The kernel that provides `amdxdna` and `amdgpu`, pinned to upstream | `torvalds/linux` release tags; config from the Strix Halo kernel of 2026-09-23 | **pinned** ([docs/kernel.md](kernel.md)): v7.3-rc4 builds into Debian packages with `amdxdna` in-tree; kept current by `bump-linux.yml`. Installing it on Strix Halo is a separate, deliberate step |