chore(train): engine-10 continued: the PQ2_0 tensor-core launches take their tiles from a counter - #75
Merged
Conversation
…r own from a counter Each block of mmvq_pq2_mma (one an SM) owned every gridDim.x-th tile, so a launch ended when its slowest block did: the blocks the DRAM served later finished a tile or so after the rest, and the launch's last microseconds ran on a few SMs. Now a block's first tiles, as many as its ring holds before the dependency wait, are still its own; past them its producer takes the next tile by an atomic add on its stream's counter (ggml_cuda_pq2_tile_counters, one int a stream, made zeroed before a graph evaluation), after its own dependency wait, when every launch before it on the stream has ended. Every block's last ticket is past the tiles, so a launch takes n_tiles - dyn_base + gridDim.x of them, and the block that takes the last sets the counter back to 0 for the next launch (CUDA graph replays included). The producer tells the consumers each slot's tile (or the end) in shared memory beside the slot's barriers. The arithmetic and its order are unchanged. GGML_CUDA_PQ2_MMA_TILES_LEGACY=1 gives every tile to its block again. RTX 5080, Ternary Bonsai 2 27B, q4_0 K/V, f16 state, one build with and without the switch: - llama-bench with graphs on, legs N-L-L-N: tg128 at depth 0 +1.46 % (106.18 / 105.74 / 106.04 / 106.07 against 103.38 / 104.83 / 105.01 / 104.69 tok/s), at 16,384 +1.34 %, tg64 at 245,760 +1.47 % (68.19 / 68.26 against 67.92 / 66.55); pp4 with -rs 3 (the MTP verify), the median of 30 samples a leg, +0.62 % at 0 and +0.81 % at 16,384. - llama-server as served (MTP drafts of 3, one slot, 32,768 context), 12 greedy agentic requests a leg: the same text in every pair, 231.53 -> 234.18 tok/s against the undisturbed legacy leg (same-text geomean +1.8 %, se 0.5 %). - nsys, graphs off, 16 tokens at 16,384: 169,594 -> 166,576 us. - test-backend-ops MUL_MAT (pq2_0) 99/99, MUL_MAT_GROUP 24/24, MUL_MAT_VEC_FUSION 1010/1010, both ways; without the last ticket's reset the counter build fails 29, 7 and 26 of them and passes them all under the switch. llama-server at one slot and at 4 slots with --kv-unified, 128 greedy tokens with top-5 log-probabilities: bit-identical to the published engine-32e695e build both ways.
Stage 1 of GLM-5.3 tensor parallelism: the split callback mirrors every KDA and DSA weight, all ssm and indexer tensors and all four caches, and keeps the existing splits for dense experts, routed experts, shared experts (as DeepSeek4) and the head, with one all-reduce per layer past the expert adds. glm5next leaves the -sm tensor refusal list. Mirrored mapping also dodges the ssm_out suffix assert and the ssm_d_state granularity divide for this arch. test-llama-archs -a glm5next, local 5080 + 5070 Ti: Meta row OK at 2.82e-06 NMSE vs CPU (1e-4 bar). Mutation proof: ffn_down_exps mirrored while up stays split aborts the Meta row in the matmul split-state handler, so the green row really splits. Per-card weights from the GGUF tensor table: 66.15 GiB plus mirrored latent and indexer caches.
…s 512 pools ggml_cuda_op_top_k took a k of 512 over a few thousand columns (the DSA indexer's pools, one launch a layer) through the tiled path, whose one block a tile runs k block-wide max reductions one after another: 194 us a call at 3,520 columns. topk_radix finds the k-th largest key in four 8-bit histogram passes (warp-aggregated shared atomics, one per digit a warp), then writes in column order every column above it and the lowest columns equal to it, so the set and its tie-break are the tiled path's. A row of up to 16,384 columns keeps its keys in registers; a wider one is cut into 8,192- column tiles whose blocks select their k in parallel (a key among the row's k best has fewer than k better keys, so its tile keeps it), and one block a row selects among the tiles' candidates. One wide row stays with CUB's device-wide top-k, which spreads it over the card. Taken for k >= 64 past 1,024 columns on NVIDIA from Volta; GGML_CUDA_TOPK_RADIX_LEGACY=1 keeps the old dispatch. RTX 5080, test-backend-ops perf, legs new-legacy-new-legacy, k 512: 3,520 columns 194.5 -> 6.1 us (1 and 3 rows), 8,192 x 3 199.3 -> 12.4, 32,768 x 3 58.3 -> 17.1, 109,020 x 3 66.3 -> 21.8, and a 512-row ubatch 13.8-18x at every width; one wide row unchanged (CUB both ways). TOP_K 523/523 (new cases at the indexer's widths with and without ties). Two mutants fail it: the last equal key never written (every radix case), and stage 2 writing a candidate's position for its column (the 19 two-stage cases, GLM's 109,020 x 3 among them, and no other). The tests also gain GLM-5.3-Flash perf shapes: its routed experts (288, 8 used, IQ3_XXS / IQ4_XS / Q8_0) at a decode, an MTP verify and a ubatch, and its dense Q8_0 projections. Seen on the legacy dispatch, not fixed here: one of five legacy perf runs aborted at ggml-cuda.cu:3058 (cudaGraphExecUpdate returning neither success nor an update failure) on the graph after CUB's per-row top-k at 151,936 x 16, k 64. The radix path does not take CUB for more than one row.
…x of the cells build_attn_mha tagged every one-sequence causal mask as a prefix, the hint under which the CUDA flash attention leaves out its range scan and live-tile skipping and applies the mask over the whole range (fattn-common.cuh, mask_prefix). The DSA layers (GLM-5.3's MLA with the lightning indexer, and the kpool path) pass a mask that keeps the indexer's top-k cells of the whole cache, so a decode read every cell's K and V to keep 2,048 of them. build_attn_mha takes mask_is_prefix, true by default; both sparse callers pass false, and the kernel finds their live tiles by scanning the mask. The dense layers keep the hint. LLAMA_ATTN_SPARSE_MASK_PREFIX_LEGACY=1 tags the sparse masks again. The result is the same attention by a different split of the cells (stream-k over the live steps), so its bits can move at float rounding. No other path changes: the no-hint route is what every masked decode without the hint already takes, and test-backend-ops' FLASH_ATTN_EXT covers the hint both ways at 576/512. Checked here only that the tiny glm5next model decodes at 6,000 cells both ways; the gain sits where the cache is long against the top-k (a 436k-cell GLM-5.3 decode), measured on the served model.
…ow after another A row of up to 1,024 columns has no tiled path, so with CUB's top-k available ggml_cuda_op_top_k looped over the rows, four launches a row: a 512-token ubatch through a DSA indexer under 4,096 cells (its pools at most 1,024) cost 2,048 launches a layer (seen under nsys on the tiny glm5next model's depth prefill). topk_radix takes those rows at any k, one block a row; a wide row with a small k keeps the tiled path. RTX 5080, test-backend-ops perf, legs new-legacy-new-legacy: a narrow row 3.3-3.9 us against 8.2-12.7 (1 row) and 129-171 (16 rows), every k from 1 to 400. TOP_K 523/523, its narrow cases (ties included) now through the radix kernel that the two mutants of 881c823 fail. The GLM-5.3 routed-expert perf cases build 32 experts of the model's 288: the case quantizes every expert on the CPU at setup, and all 288 took over ten minutes; 32 are still past a 5080's L2.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The shared branch's commits after #74:
GGML_CUDA_PQ2_MMA_TILES_LEGACY=1restores the old assignment. Measured on an RTX 5080 (Ternary Bonsai 2 27B): tg +1.3 to +1.5 %, and the served MTP decode goes from 231.53 to 234.18 tok/s with identical text.