Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions docs/hrx.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,34 @@ Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan):

`test-backend-ops -b HRX0`: 790/790.

### Fixed: decode-split multipass gap above 2048 ([engine#124](https://github.com/1bit-MONSTER/engine/issues/124))

The multipass decode-split reducer (`reduce_completed.multipass`, used above
`key_value_token_capacity` 2048) finished with a scalar output pass: each workitem owned
one output channel and walked every KV block serially, recomputing
`expf(partial_max - maximum)` once per (block, element). Only half the workgroup is live
at `value_head_size = 128`, so decode above 2048 stayed about 16% below the <=2048
corridor. Fork PR
[1bit-MONSTER/llama.cpp#30](https://github.com/1bit-MONSTER/llama.cpp/pull/30) computes
the per-block scale once into a per-row workgroup (LDS) stage and runs one vectorised,
all-rows output pass (4 channels per workitem). The per-channel block order is
unchanged, so the reduce is bit-identical.

Measured on Strix Halo (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, `-dev HRX0`,
`llama-bench -p 0 -n 8 -r 5`, median of 6 interleaved runs):

| depth | blocks | before | this pin | delta |
|---|---:|---:|---:|---:|
| 2000 (<=2048 control) | 32 | 66.06 | 68.82 | +4.2% |
| 2100 | 33 | 58.19 | 67.15 | **+15.4%** |
| 3000 | 47 | 50.84 | 60.89 | **+19.8%** |
| 4800 | 76 | 41.83 | 51.02 | **+22.0%** |

`d2100/d2000` moves 0.881 -> 0.976, so the boundary cliff is gone. 1.24 GB of
decode-path logits (`llama-perplexity -c 2049 -b 1 --save-all-logits`) are byte-identical
to the previous pin, the buried code word is exact at 4700 tokens, and no GPU faults
were observed.

### Known issues on `HRX0`

- **Several sequences per batch fail.** `llama-perplexity` with `n_seq` > 1 stops on an
Expand Down
2 changes: 1 addition & 1 deletion registry/architectures.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
"llama.cpp (vulkan)": "cecf3ee01d9d99378e98bfea95f51cac714b04b8",
"llama.cpp (hrx)": "00adc2b9d6676f39cb3fcca8b5ad97621e653fae",
"llama.cpp (hrx)": "fcd83eb2a4a86cdae5d65e114981470593a190aa",
"zinc": "29bc350ac4cbf9f110ec628b9e177ea04ac816fb"
},
"counts": {
Expand Down
Loading