Skip to content

chore(train): engine-10 continued: llama-bench's generation range, the rotations' weights before the PDL wait - #74

Merged
marcospaulo merged 4 commits into
mainfrom
train/engine-10
Sep 27, 2026
Merged

marcospaulo merged 4 commits into
mainfrom
train/engine-10

Conversation

@marcospaulo

@marcospaulo marcospaulo commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

The train engine-10 past #73: 4 commits (e75e318, 017848c, fb03d29, 7a62668).

llama-bench's generation range (e75e318). An NVTX range gen (domain llama-bench, a registered string) wraps each repetition's generation. nsys profile -c nvtx -p gen@llama-bench --capture-range-end=stop then records the tokens and none of the depth's prefill.

At a depth of 245,760 on an RTX 5080, the full capture's ~1.19M prefill kernels kept nsys past rig's 30-minute probe limit. Started at the range it takes 179 s, a 2.3 MB report. At 16,384 it reads what the full capture reads: kernels 9,553.6 against 9,559.8 us a token, the same 23 groups.

The rotations' weights before the PDL wait (fb03d29). rms_norm_fwht_cuda and fwht_cuda_block read their norm weight and signs after the grid dependency wait. The PQ2_0 matmul after a rotation starts under it, so its ring fill and the L2 prefetch of its head took DRAM while those weights queued behind them. They are never written by a kernel, so they are now requested before the wait (the SASS has both LDGs ahead of ACQBULK).

RTX 5080, Ternary Bonsai 2 27B, one build with and without GGML_CUDA_FWHT_PREWAIT_LEGACY=1:

  • The FFN-input norm now ends 4.7-4.8 us past its out-projection instead of 6.1-6.2.
  • llama-bench, graphs on, legs N-L-L-N: +0.81 % tg128 at 16,384 and +1.63 % tg64 at 245,760, all four pairs faster.
  • test-backend-ops MUL_MAT_HADAMARD passes 42/42 both ways.
  • llama-server at one slot and at 4 slots is bit-identical to the published engine-32e695e (128 greedy tokens, top-5 log-probabilities).

…capture can start at the tokens

nsys profile -c nvtx -p gen@llama-bench --capture-range-end=stop records the first repetition's
generation and none of the depth's prefill before it. rig's roofline probe traced the whole run: at
a depth of 245,760 on an RTX 5080 (Ternary Bonsai 2 27B, CUDA graphs off, GGML_CUDA_NVTX=1) the
prefill's ~1.19M kernels (2,487 a 512-token ubatch) kept nsys past the probe's 30-minute limit.
Started at the range, the same capture takes 179 s with the prefill untraced, a 2.3 MB report of
the 16 tokens' 18,816 kernels. At 16,384 it reads what the full capture reads: 15 tokens, kernels
9,553.6 against 9,559.8 us a token, the same 23 groups, the largest group apart by 6.3 us.

The message is a registered string in the domain llama-bench, which nsys matches without
NSYS_NVTX_PROFILER_REGISTER_ONLY=0. NVTX is header-only: a CUDA build links CUDA::nvtx3, and a
build without the headers compiles the range out.
… dependency wait

rms_norm_fwht_cuda read its norm weight and signs, and fwht_cuda_block its signs, after
cudaGridDependencySynchronize. The PQ2_0 matmul after a rotation starts under it (PDL): its ring
fill and the L2 prefetch of its head (ggml_cuda_pq2_prefetch) take DRAM for the microseconds the
rotation runs, and the rotation's weights, DRAM misses each token, queued behind them. Those weights
are never written by a kernel, so, like the matmul's own, they are now requested before the wait
and arrive while the kernel before the rotation ends (SASS: the two LDGs ahead of ACQBULK). The
arithmetic and its order are unchanged.

RTX 5080, Ternary Bonsai 2 27B, q4_0 K/V, f16 state, one build with and without
GGML_CUDA_FWHT_PREWAIT_LEGACY=1:
- nsys over 15 tokens (graphs off, the capture started at llama-bench's gen range): the FFN-input
  norm after an out-projection ends 6.1-6.2 -> 4.7-4.8 us past it at 16,384 (5.5-6.1 -> 4.5-4.8
  at 245,760), the layer-input norm after ffn_down 2.1 -> 1.7 us; the median token 9,779.9 ->
  9,611.5 us at 16,384, 14,704 -> 14,624 us at 245,760.
- llama-bench with graphs on, legs N-L-L-N: tg128 at 16,384 105.48 / 104.21 / 104.40 / 104.82
  tok/s (+0.81 %, both pairs faster); tg64 at 245,760 69.35 / 67.93 / 67.48 / 68.27 (+1.63 %,
  both pairs faster).
- test-backend-ops MUL_MAT_HADAMARD 42/42 both ways; llama-server at one slot and at 4 slots with
  --kv-unified, 128 greedy tokens with top-5 log-probabilities on a 15,616-token prompt:
  bit-identical to the published engine-32e695e build both ways.
@marcospaulo marcospaulo changed the title chore(train): engine-10 continued: llama-bench's generation range, so a profile captures the tokens chore(train): engine-10 continued: llama-bench's generation range, the rotations' weights before the PDL wait Sep 27, 2026
@marcospaulo
marcospaulo merged commit 4f13c9c into main Sep 27, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant