Skip to content

chore(train): engine-10 continued: the PQ2_0 tensor-core launches take their tiles from a counter - #75

Merged
marcospaulo merged 10 commits into
mainfrom
train/engine-10
Sep 28, 2026
Merged

marcospaulo merged 10 commits into
mainfrom
train/engine-10

Conversation

@marcospaulo

Copy link
Copy Markdown
Member

The shared branch's commits after #74:

  • 8c78ba9 perf(cuda): each PQ2_0 tensor-core launch hands out tiles from a per-stream counter, so a launch no longer waits on its slowest block. The arithmetic and its order are unchanged. GGML_CUDA_PQ2_MMA_TILES_LEGACY=1 restores the old assignment. Measured on an RTX 5080 (Ternary Bonsai 2 27B): tg +1.3 to +1.5 %, and the served MTP decode goes from 231.53 to 234.18 tok/s with identical text.
  • a28b768 docs(torad): the TORAD.md row for it.

…r own from a counter

Each block of mmvq_pq2_mma (one an SM) owned every gridDim.x-th tile, so a launch ended when its
slowest block did: the blocks the DRAM served later finished a tile or so after the rest, and the
launch's last microseconds ran on a few SMs. Now a block's first tiles, as many as its ring holds
before the dependency wait, are still its own; past them its producer takes the next tile by an
atomic add on its stream's counter (ggml_cuda_pq2_tile_counters, one int a stream, made zeroed
before a graph evaluation), after its own dependency wait, when every launch before it on the
stream has ended. Every block's last ticket is past the tiles, so a launch takes
n_tiles - dyn_base + gridDim.x of them, and the block that takes the last sets the counter back
to 0 for the next launch (CUDA graph replays included). The producer tells the consumers each
slot's tile (or the end) in shared memory beside the slot's barriers. The arithmetic and its
order are unchanged. GGML_CUDA_PQ2_MMA_TILES_LEGACY=1 gives every tile to its block again.

RTX 5080, Ternary Bonsai 2 27B, q4_0 K/V, f16 state, one build with and without the switch:
- llama-bench with graphs on, legs N-L-L-N: tg128 at depth 0 +1.46 % (106.18 / 105.74 / 106.04 /
  106.07 against 103.38 / 104.83 / 105.01 / 104.69 tok/s), at 16,384 +1.34 %, tg64 at 245,760
  +1.47 % (68.19 / 68.26 against 67.92 / 66.55); pp4 with -rs 3 (the MTP verify), the median
  of 30 samples a leg, +0.62 % at 0 and +0.81 % at 16,384.
- llama-server as served (MTP drafts of 3, one slot, 32,768 context), 12 greedy agentic
  requests a leg: the same text in every pair, 231.53 -> 234.18 tok/s against the undisturbed
  legacy leg (same-text geomean +1.8 %, se 0.5 %).
- nsys, graphs off, 16 tokens at 16,384: 169,594 -> 166,576 us.
- test-backend-ops MUL_MAT (pq2_0) 99/99, MUL_MAT_GROUP 24/24, MUL_MAT_VEC_FUSION 1010/1010,
  both ways; without the last ticket's reset the counter build fails 29, 7 and 26 of them and
  passes them all under the switch. llama-server at one slot and at 4 slots with --kv-unified,
  128 greedy tokens with top-5 log-probabilities: bit-identical to the published engine-32e695e
  build both ways.
Stage 1 of GLM-5.3 tensor parallelism: the split callback mirrors every
KDA and DSA weight, all ssm and indexer tensors and all four caches, and
keeps the existing splits for dense experts, routed experts, shared
experts (as DeepSeek4) and the head, with one all-reduce per layer past
the expert adds. glm5next leaves the -sm tensor refusal list. Mirrored
mapping also dodges the ssm_out suffix assert and the ssm_d_state
granularity divide for this arch.

test-llama-archs -a glm5next, local 5080 + 5070 Ti: Meta row OK at
2.82e-06 NMSE vs CPU (1e-4 bar). Mutation proof: ffn_down_exps mirrored
while up stays split aborts the Meta row in the matmul split-state
handler, so the green row really splits. Per-card weights from the GGUF
tensor table: 66.15 GiB plus mirrored latent and indexer caches.
…s 512 pools

ggml_cuda_op_top_k took a k of 512 over a few thousand columns (the DSA indexer's pools, one launch a
layer) through the tiled path, whose one block a tile runs k block-wide max reductions one after
another: 194 us a call at 3,520 columns. topk_radix finds the k-th largest key in four 8-bit
histogram passes (warp-aggregated shared atomics, one per digit a warp), then writes in column order
every column above it and the lowest columns equal to it, so the set and its tie-break are the tiled
path's. A row of up to 16,384 columns keeps its keys in registers; a wider one is cut into 8,192-
column tiles whose blocks select their k in parallel (a key among the row's k best has fewer than k
better keys, so its tile keeps it), and one block a row selects among the tiles' candidates. One
wide row stays with CUB's device-wide top-k, which spreads it over the card. Taken for k >= 64 past
1,024 columns on NVIDIA from Volta; GGML_CUDA_TOPK_RADIX_LEGACY=1 keeps the old dispatch.

RTX 5080, test-backend-ops perf, legs new-legacy-new-legacy, k 512: 3,520 columns 194.5 -> 6.1 us
(1 and 3 rows), 8,192 x 3 199.3 -> 12.4, 32,768 x 3 58.3 -> 17.1, 109,020 x 3 66.3 -> 21.8, and a
512-row ubatch 13.8-18x at every width; one wide row unchanged (CUB both ways). TOP_K 523/523
(new cases at the indexer's widths with and without ties). Two mutants fail it: the last equal key
never written (every radix case), and stage 2 writing a candidate's position for its column (the 19
two-stage cases, GLM's 109,020 x 3 among them, and no other).

The tests also gain GLM-5.3-Flash perf shapes: its routed experts (288, 8 used, IQ3_XXS / IQ4_XS /
Q8_0) at a decode, an MTP verify and a ubatch, and its dense Q8_0 projections.

Seen on the legacy dispatch, not fixed here: one of five legacy perf runs aborted at ggml-cuda.cu:3058
(cudaGraphExecUpdate returning neither success nor an update failure) on the graph after CUB's
per-row top-k at 151,936 x 16, k 64. The radix path does not take CUB for more than one row.
…x of the cells

build_attn_mha tagged every one-sequence causal mask as a prefix, the hint under which the CUDA flash
attention leaves out its range scan and live-tile skipping and applies the mask over the whole range
(fattn-common.cuh, mask_prefix). The DSA layers (GLM-5.3's MLA with the lightning indexer, and the
kpool path) pass a mask that keeps the indexer's top-k cells of the whole cache, so a decode read every
cell's K and V to keep 2,048 of them. build_attn_mha takes mask_is_prefix, true by default; both sparse
callers pass false, and the kernel finds their live tiles by scanning the mask. The dense layers keep
the hint. LLAMA_ATTN_SPARSE_MASK_PREFIX_LEGACY=1 tags the sparse masks again.

The result is the same attention by a different split of the cells (stream-k over the live steps), so
its bits can move at float rounding. No other path changes: the no-hint route is what every masked
decode without the hint already takes, and test-backend-ops' FLASH_ATTN_EXT covers the hint both ways
at 576/512. Checked here only that the tiny glm5next model decodes at 6,000 cells both ways; the gain
sits where the cache is long against the top-k (a 436k-cell GLM-5.3 decode), measured on the served
model.
…ow after another

A row of up to 1,024 columns has no tiled path, so with CUB's top-k available ggml_cuda_op_top_k looped
over the rows, four launches a row: a 512-token ubatch through a DSA indexer under 4,096 cells (its
pools at most 1,024) cost 2,048 launches a layer (seen under nsys on the tiny glm5next model's depth
prefill). topk_radix takes those rows at any k, one block a row; a wide row with a small k keeps the
tiled path.

RTX 5080, test-backend-ops perf, legs new-legacy-new-legacy: a narrow row 3.3-3.9 us against 8.2-12.7
(1 row) and 129-171 (16 rows), every k from 1 to 400. TOP_K 523/523, its narrow cases (ties included)
now through the radix kernel that the two mutants of 881c823 fail.

The GLM-5.3 routed-expert perf cases build 32 experts of the model's 288: the case quantizes every
expert on the CPU at setup, and all 288 took over ten minutes; 32 are still past a 5080's L2.
@marcospaulo
marcospaulo merged commit 9628ba2 into main Sep 28, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant