chore(train): engine-10 continued: llama-bench's generation range, the rotations' weights before the PDL wait - #74
Merged
Merged
Conversation
…capture can start at the tokens nsys profile -c nvtx -p gen@llama-bench --capture-range-end=stop records the first repetition's generation and none of the depth's prefill before it. rig's roofline probe traced the whole run: at a depth of 245,760 on an RTX 5080 (Ternary Bonsai 2 27B, CUDA graphs off, GGML_CUDA_NVTX=1) the prefill's ~1.19M kernels (2,487 a 512-token ubatch) kept nsys past the probe's 30-minute limit. Started at the range, the same capture takes 179 s with the prefill untraced, a 2.3 MB report of the 16 tokens' 18,816 kernels. At 16,384 it reads what the full capture reads: 15 tokens, kernels 9,553.6 against 9,559.8 us a token, the same 23 groups, the largest group apart by 6.3 us. The message is a registered string in the domain llama-bench, which nsys matches without NSYS_NVTX_PROFILER_REGISTER_ONLY=0. NVTX is header-only: a CUDA build links CUDA::nvtx3, and a build without the headers compiles the range out.
… dependency wait rms_norm_fwht_cuda read its norm weight and signs, and fwht_cuda_block its signs, after cudaGridDependencySynchronize. The PQ2_0 matmul after a rotation starts under it (PDL): its ring fill and the L2 prefetch of its head (ggml_cuda_pq2_prefetch) take DRAM for the microseconds the rotation runs, and the rotation's weights, DRAM misses each token, queued behind them. Those weights are never written by a kernel, so, like the matmul's own, they are now requested before the wait and arrive while the kernel before the rotation ends (SASS: the two LDGs ahead of ACQBULK). The arithmetic and its order are unchanged. RTX 5080, Ternary Bonsai 2 27B, q4_0 K/V, f16 state, one build with and without GGML_CUDA_FWHT_PREWAIT_LEGACY=1: - nsys over 15 tokens (graphs off, the capture started at llama-bench's gen range): the FFN-input norm after an out-projection ends 6.1-6.2 -> 4.7-4.8 us past it at 16,384 (5.5-6.1 -> 4.5-4.8 at 245,760), the layer-input norm after ffn_down 2.1 -> 1.7 us; the median token 9,779.9 -> 9,611.5 us at 16,384, 14,704 -> 14,624 us at 245,760. - llama-bench with graphs on, legs N-L-L-N: tg128 at 16,384 105.48 / 104.21 / 104.40 / 104.82 tok/s (+0.81 %, both pairs faster); tg64 at 245,760 69.35 / 67.93 / 67.48 / 68.27 (+1.63 %, both pairs faster). - test-backend-ops MUL_MAT_HADAMARD 42/42 both ways; llama-server at one slot and at 4 slots with --kv-unified, 128 greedy tokens with top-5 log-probabilities on a 15,616-token prompt: bit-identical to the published engine-32e695e build both ways.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The train engine-10 past #73: 4 commits (e75e318, 017848c, fb03d29, 7a62668).
llama-bench's generation range (e75e318). An NVTX range
gen(domainllama-bench, a registered string) wraps each repetition's generation.nsys profile -c nvtx -p gen@llama-bench --capture-range-end=stopthen records the tokens and none of the depth's prefill.At a depth of 245,760 on an RTX 5080, the full capture's ~1.19M prefill kernels kept nsys past rig's 30-minute probe limit. Started at the range it takes 179 s, a 2.3 MB report. At 16,384 it reads what the full capture reads: kernels 9,553.6 against 9,559.8 us a token, the same 23 groups.
The rotations' weights before the PDL wait (fb03d29).
rms_norm_fwht_cudaandfwht_cuda_blockread their norm weight and signs after the grid dependency wait. The PQ2_0 matmul after a rotation starts under it, so its ring fill and the L2 prefetch of its head took DRAM while those weights queued behind them. They are never written by a kernel, so they are now requested before the wait (the SASS has both LDGs ahead of ACQBULK).RTX 5080, Ternary Bonsai 2 27B, one build with and without
GGML_CUDA_FWHT_PREWAIT_LEGACY=1:test-backend-ops MUL_MAT_HADAMARDpasses 42/42 both ways.