Skip to content

Pin llama.cpp a34a6b7: HRX decode-split multipass output pass is vectorised (engine#124) - #180

Merged
bong-water-water-bong merged 3 commits into
mainfrom
hrx-124-multipass-output
Sep 27, 2026
Merged

bong-water-water-bong merged 3 commits into
mainfrom
hrx-124-multipass-output

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Pins 1bit-MONSTER/llama.cpp#30 and documents it. Closes #124.

What

The multipass decode-split reducer (reduce_completed.multipass, used above key_value_token_capacity 2048) ended with a scalar output pass: each workitem owned one output channel and walked every KV block serially, recomputing expf(partial_max - maximum) once per (block, element). Only half the workgroup is live at value_head_size = 128, so decode above 2048 stayed ~16% below the <=2048 corridor.

The per-block scale is now computed once in the lane-strided sum pass into a per-row workgroup (LDS) stage, and one vectorised all-rows pass (4 channels per workitem) does the output. The per-channel block order is unchanged, so the reduce is bit-identical.

Measured (Strix Halo, Qwen3-Coder-30B-A3B-Instruct Q4_K_M, -dev HRX0)

llama-bench -p 0 -n 8 -r 5, median of 6 interleaved runs:

depth blocks base patched delta
1900 30 70.72 71.48 +1.1% (control)
2000 32 66.06 68.82 +4.2% (control)
2100 33 58.19 67.15 +15.4%
3000 47 50.84 60.89 +19.8%
4800 76 41.83 51.02 +22.0%

d2100/d2000 moves 0.881 -> 0.976 — the boundary cliff is gone. Nothing changes at or below capacity 2048.

Evidence

  • Bit-exact: 1,244,725,284-byte decode-path logits dumps (llama-perplexity -c 2049 -b 1 --save-all-logits; GGML_HRX_LOG_DISPATCH=1 confirms flash_attention_decode_split) are byte-identical to the previous pin; PPL 21.9760 +/- 0.81813 on both.
  • Correct: buried code word ZX-4718-QQ exact at 4700 prompt tokens (capacity 4864, multipass), temperature 0, seed 42, cache_prompt:false.
  • No faults: 6 rounds x 5-rep sweeps at all five depths, 0 HSA_STATUS_ERROR_MEMORY_FAULT, 0 res = -3.

Full table, method and repro: third_party/llama.cpp -> benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md.

Docs: new docs/hrx.md "Fixed: decode-split multipass gap above 2048 (engine#124)" section.

Pin updated: llama.cpp #30 is merged, so this now pins the branch head fcd83eb (a merge commit with the same tree as a34a6b7); registry/architectures.json regenerated for it, and main merged in.

Review (strixhalo, fork build of a34a6b7 vs its base 00adc2b, same toolchain)

  • Kernel: the new scale and sum stages are LDS-only and fenced by barrier<workgroup>; the pack's barrier<global> + barrier<workgroup> from Pin llama.cpp 00adc2b: HRX decode-split publishes next_q8 after a global barrier (#123, #140) #176 is untouched.
  • HRX: decode-split multipass path intermittently faults (HSA_STATUS_ERROR_MEMORY_FAULT) above capacity 2048 #123 regression, multipass path: Qwen3-Coder-30B, 2113-token prompt, GGML_HRX_FA_PARTIAL_ALIGN=256, a separate process forcing KFD queue evictions: 8/8 correct, 0 NaN.
  • Bit-exact: decode-path perplexity (llama-perplexity -c 2304 -b 1 --chunks 1, Coder-30B) = 7.5538 +/- 0.66265 on both builds.
  • Speed (llama-bench -p 0 -n 16 -r 3, 2 interleaved rounds, tok/s): d2000 69.7/68.7 -> 67.5/67.3 (noise), d2100 48.9/49.7 -> 68.0/64.1, d4800 40.2/41.7 -> 42.4/47.0.
  • LDS: the new scale_stage grows with the block count. JIT-compiled at capacity 32768: 34.1 KB (32/4 heads) and 50.5 KB (64/4) vs a flat 17.7 KB before, under the 64 KB limit. Both builds fail to compile the kernel above 32768 (40960 to 262144 tried), so that ceiling is pre-existing and this PR does not move it.

…orised (engine#124)

Brings 1bit-MONSTER/llama.cpp#30. The multipass reducer's output pass gave each
workitem one channel and walked every KV block serially, recomputing
expf(partial_max - maximum) per (block, element); only half the workgroup is live
at value_head_size=128, which was the whole residual >2048 decode gap.

The per-block scale is now computed once in the lane-strided sum pass into a
per-row LDS stage, and one vectorised all-rows pass (4 channels per workitem) does
the output. Per-channel block order is unchanged, so the reduce is bit-identical
(1.24 GB of decode-path logits compare equal to the previous pin).

d2100 +15.4%, d3000 +19.8%, d4800 +22.0%; d2100/d2000 0.881 -> 0.976, so the
boundary cliff is gone. Nothing changes at or below capacity 2048. Details and
repro in benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md.
@context7

context7 Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 20181e4

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 20181e4)

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Performance Claim Verification

The PR claims a 22% performance improvement for decode above 2048 tokens, measured on Strix Halo with Qwen3-Coder-30B-A3B-Instruct Q4_K_M. However, the PR description does not provide specific details on how the performance gains were validated or whether the measurements were conducted under controlled conditions. This lack of detail makes it difficult to independently verify the claimed performance improvements.

### Fixed: decode-split multipass gap above 2048 ([engine#124](https://github.com/1bit-MONSTER/engine/issues/124))

The multipass decode-split reducer (`reduce_completed.multipass`, used above
`key_value_token_capacity` 2048) finished with a scalar output pass: each workitem owned
one output channel and walked every KV block serially, recomputing
`expf(partial_max - maximum)` once per (block, element). Only half the workgroup is live
at `value_head_size = 128`, so decode above 2048 stayed about 16% below the <=2048
corridor. Fork PR
[1bit-MONSTER/llama.cpp#30](https://github.com/1bit-MONSTER/llama.cpp/pull/30) computes
the per-block scale once into a per-row workgroup (LDS) stage and runs one vectorised,
all-rows output pass (4 channels per workitem). The per-channel block order is
unchanged, so the reduce is bit-identical.

Measured on Strix Halo (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, `-dev HRX0`,
`llama-bench -p 0 -n 8 -r 5`, median of 6 interleaved runs):

| depth | blocks | before | this pin | delta |
|---|---:|---:|---:|---:|
| 2000 (<=2048 control) | 32 | 66.06 | 68.82 | +4.2% |
| 2100 | 33 | 58.19 | 67.15 | **+15.4%** |
| 3000 | 47 | 50.84 | 60.89 | **+19.8%** |
| 4800 | 76 | 41.83 | 51.02 | **+22.0%** |

`d2100/d2000` moves 0.881 -> 0.976, so the boundary cliff is gone. 1.24 GB of
decode-path logits (`llama-perplexity -c 2049 -b 1 --save-all-logits`) are byte-identical
to the previous pin, the buried code word is exact at 4700 tokens, and no GPU faults
were observed.

@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 1e4847f

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Blocked on 1bit-MONSTER/llama.cpp#30: the new workgroup scale stage grows with context length (rows × blocks × 4 B) and can pass 64 KB of LDS at 64K-131K tokens, while the PR was measured only up to d4800. Details: 1bit-MONSTER/llama.cpp#30 (comment). Re-pin here once #30 bounds it or shows long-context runs are safe.

bong-water-water-bong and others added 2 commits September 27, 2026 19:14
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 20181e4

@bong-water-water-bong
bong-water-water-bong merged commit 56f6439 into main Sep 27, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-124-multipass-output branch September 27, 2026 22:16
bong-water-water-bong added a commit that referenced this pull request Sep 28, 2026
…and PORTING status (#182)

The decode-split q8 pack read global output behind an LDS-only barrier; fixed in
llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised
(#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the
NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

HRX: decode-split multipass reduce output pass is O(blocks) per element, leaving a ~16% decode gap above 2048

1 participant