Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ Windows and 1bit OS ([docs/releases.md](docs/releases.md)).
> [1bit-MONSTER/lemonade](https://github.com/1bit-MONSTER/lemonade)) passes Lemonade's own LLM test
> suite on Vulkan and HRX. Following geramyL's review the engine no longer
> vendors Lemonade: Lemonade is the host, 1bit is the engine inside it. Ported so far: HRX on AMD's live ggml-hrx
> ([docs/hrx.md](docs/hrx.md)), Vulkan from upstream llama.cpp's latest release ([docs/vulkan.md](docs/vulkan.md)), the NPU engine on full ELFs with the
> ([docs/hrx.md](docs/hrx.md); its decode-split race, #123/#140, is fixed and the kernel is on by default), Vulkan from upstream llama.cpp's latest release ([docs/vulkan.md](docs/vulkan.md)), the NPU engine on full ELFs with the
> upstream XDNA stack pinned ([docs/npu.md](docs/npu.md); its layer kernel is not yet built from
> source), ZINC ([docs/zinc.md](docs/zinc.md)) and MLX ([docs/apple.md](docs/apple.md)). ZAYA1-8B (Zyphra)
> runs from our llama.cpp on Vulkan, HRX and ROCm, matching transformers, at 93 tok/s decode in
Expand Down
80 changes: 80 additions & 0 deletions blog/2026-09-27-hrx-decode-race-solved.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
<!--
Copyright 2026 bong-water-water-bong
SPDX-License-Identifier: Apache-2.0

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
-->
tags: hrx, milestone, kernels

# Milestone: the HRX decode race is solved

Yesterday's post ended with a workaround. HRX's split flash-attention decode kernel gave
different answers to identical requests, and on Qwen3-Coder-30B it sometimes produced NaNs that
faulted the GPU ([#123](https://github.com/1bit-MONSTER/engine/issues/123),
[#140](https://github.com/1bit-MONSTER/engine/issues/140)). So `1bit serve --device hrx` stopped
using it. Today we found the cause and fixed it, and the kernel is back on by default. It is
correct again, and faster than the path we used without it.

## What it was

The decode-split kernel attends over the KV cache in blocks, and a final workgroup combines the
blocks. That workgroup writes the attention output to global memory. Other waves then read it
back to quantize it (`next_q8`) for the next projection. Between the write and the read sat a
workgroup barrier that fences shared memory (LDS) only. It does not wait for global stores or
invalidate the cache the readers use. So a reading wave could pick up the output slot's
*previous* contents. The f32 output was right, the quantized copy was wrong, and when the stale
bytes happened to decode as NaN the MoE router saw NaN logits.

The window only opens when the GPU preempts the dispatch in the middle. Any process's KFD queue
eviction does that: page compaction, KSM, memory pressure. That is why the failure rate followed
those system settings. For a while it looked like a page-migration bug, but migration was only
the trigger. The fix fences both memory spaces at all five places the kernel does this:
`barrier<global>`, then `barrier<workgroup>`
([docs/hrx.md](../docs/hrx.md#moe-router-expert-id-fault-fix-engine123)).

## How we know

A separate process forced queue evictions while the engine served, to make the rare case common:

- **Qwen3-Coder-30B-A3B** (2113-token prompt): 1 of 8 correct before, 8 of 8 after, with no NaN.
- **Qwen3-0.6B**, greedy: 13-34 divergent answers per ~177 requests before, none in 354 after.

The last clue came from hashing every kernel's output in serial execution. In every bad request,
the first command whose output diverged was this kernel, with the f32 output identical and the
quantized output different.

## Faster with it on

Decode tok/s (Q4_K_M, llama-bench, interleaved on a shared box):

| model | ctx 2100, split on | ctx 2100, split off |
|---|---|---|
| Qwen3-0.6B | 140-158 | 118-122 |
| Qwen3-Coder-30B-A3B | 42-62 | 47-49 |
| ZAYA1-8B | 38-42 | 34-36 |

The kernel's multi-pass output step was also vectorised
([#180](https://github.com/1bit-MONSTER/engine/pull/180)): +30-37% at depth 2100, with
bit-identical perplexity. `ONEBIT_HRX_DECODE_SPLIT=0` turns the kernel off if you need to compare.

## Also this week

- **Architecture gaps closed.** OPT, GPT-Neo, CodeGen and GPT-J, which upstream llama.cpp has no
model code for, and Zyphra's whole family, from ZAYA1-74B to the vision models, all run from our
llama.cpp. The registry now maps 323 Hugging Face architectures, up from 265 at step 5
([docs/registry.md](../docs/registry.md)).
- **GGUF on the NPU**: a repacked GGUF serves on the NPU. Qwen2.5-7B, MiniCPM4-8B and
MiniCPM5-1B answer there
([#179](https://github.com/1bit-MONSTER/engine/pull/179), [docs/npu.md](../docs/npu.md#from-a-gguf)).

Packages with the fix ship on Sunday with the weekly release.
2 changes: 1 addition & 1 deletion docs/PORTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ Each step below is one PR (or a short series) that builds and runs on Strix Halo
| Step | Component | Source in 1bit-MONSTER | Done when |
|---|---|---|---|
| 1 | **The engine runs inside Lemonade.** Lemonade runs `1bit serve` as a backend: one model per process behind an OpenAI-compatible API | `tools/unified_server.cpp` | **landed** ([docs/serve.md](serve.md)): NPU, Vulkan, HRX and ZINC pass `serve_e2e` on Strix Halo; `smoke_serve` runs in CI. The engine first embedded Lemonade (v11.9.0 plus local recipes); after geramyL's review it is embedded *into* Lemonade instead: the vendored copy is removed, and the Lemonade recipe that runs `1bit serve` (`onebit`) lives in our fork [1bit-MONSTER/lemonade](https://github.com/1bit-MONSTER/lemonade) ([docs/lemonade.md](lemonade.md)) |
| 2 | **HRX with Vulkan.** llama.cpp with `GGML_HRX=ON` and `GGML_VULKAN=ON` in one build, on AMD's tested pair | `1bit-MONSTER/llama.cpp` `1bit/hrx-vulkan-patched` (AMD's `hrx-graph-develop-v2` plus our commits) + `ROCm/hrx-system`, pinned from `ROCm/ggml-staging-automation` | **landed** ([docs/hrx.md](hrx.md)): `1bit serve` runs the same checkpoint on `Vulkan0` and `HRX0`; kept current by `bump-hrx.yml` |
| 2 | **HRX with Vulkan.** llama.cpp with `GGML_HRX=ON` and `GGML_VULKAN=ON` in one build, on AMD's tested pair | `1bit-MONSTER/llama.cpp` `1bit/hrx-vulkan-patched` (AMD's `hrx-graph-develop-v2` plus our commits) + `ROCm/hrx-system`, pinned from `ROCm/ggml-staging-automation` | **landed** ([docs/hrx.md](hrx.md)): `1bit serve` runs the same checkpoint on `Vulkan0` and `HRX0`; kept current by `bump-hrx.yml`. The decode-split flash-attention race (#123, #140) is fixed (llama.cpp `00adc2b`, #176) and the kernel is on by default again (#178) |
| 3 | **NPU engine.** Full ELFs only, and the open 16-tile layer kernel built from source | `engine/npu` (`npu_engine_universal.cpp`, `I8Ctx::init_elf`); ELF dispatch table on `backup/iso-build-elf-native-2026-09-22`; kernel on `bench/fastlane-16tile-corrections-2026-09-22` | **3a–3c landed** ([docs/npu.md](npu.md)). 3a (#8): full ELFs generated in C++, reproducing all 5,153 captured contexts. 3b/3c (#9): the lane runtime (logits bit-identical to the reference lane, 24/24 steps, 11.0 ms/token), tokenizer, `1bit unified`, and Qwen3-0.6B served on the NPU (now through `1bit serve`). The XDNA driver and XRT are pinned upstream and built privately (#12). **Open:** the layer kernel and lm-head artifacts are not yet built from source (npu.md, "Open") |
| 3x | **35B MoE on the NPU (experimental, closed source).** Qwen3.6-35B-A3B on the NPU | the private `1bit-MONSTER/npu-kernels` repository | **landed as a private add-on** ([docs/npu.md](npu.md#private-routes)): built into `1bit` with `-DONEBIT_NPU_PRIVATE`; parity against the fp64 reference passes (3 positions, argmax 846 / 198 / 3710), 16.3-16.5 tok/s decode, and `1bit serve --device npu` answers "The capital of France is Paris." Without the add-on, `serve` says the route is not part of the build |
| + | **Linux kernel.** The kernel that provides `amdxdna` and `amdgpu`, pinned to upstream | `torvalds/linux` release tags; config from the Strix Halo kernel of 2026-09-23 | **pinned** ([docs/kernel.md](kernel.md)): v7.3-rc4 builds into Debian packages with `amdxdna` in-tree; kept current by `bump-linux.yml`. Installing it on Strix Halo is a separate, deliberate step |
Expand Down
Loading