Skip to content

perf(cuda): cap the CUDA graphs a context keeps by their nodes, not their count - #77

Merged
marcospaulo merged 2 commits into
mainfrom
train/engine-10
Sep 28, 2026
Merged

marcospaulo merged 2 commits into
mainfrom
train/engine-10

Conversation

@marcospaulo

Copy link
Copy Markdown
Member

The shared branch's commits after #76:

  • 83f230e perf(cuda): the CUDA graphs a context keeps are now capped by their nodes, not their count.
    • A graph executable's device memory follows its nodes: 1.9-3.0 KiB a node on an RTX 5080, driver 610.43.
    • That memory comes from a driver pool that never shrinks. Destroying an executable returns nothing, and the pool stays at the most nodes held at once.
    • d0f8bae's count cap of 8 broke -sm tensor. The meta backend computes a graph per all-reduce step on each device, two a layer. Each was evicted before its reuse and captured anew every token: qwen3-0.6b decoded at 257.0 tok/s against 558.0 uncapped on a 5080 + 5070 Ti.
    • GGML_CUDA_GRAPH_NODES (default 32768, about 88 MiB) now bounds the nodes held, evicting the least recently used. GGML_CUDA_GRAPH_MAX bounds the count only when set. GGML_CUDA_GRAPH_NODES=0 GGML_CUDA_GRAPH_MAX=8 restores d0f8bae.
    • The log reports the nodes held and all the device memory the executables took.
  • f021f83 docs(torad): the TORAD.md row for it.

Checks:

  • The pool, probed directly: a cap on nodes held of 4,096, 16,384 and 32,768 left it at 14, 44 and 88 MiB through 3,500 instantiations and evictions of executables of 20-4,000 nodes.
  • test-cuda-graph-cap-nodes (new) and test-cuda-graph-cap, test-cuda-graph-key, test-cuda-graph-src-type pass. They ran in a clean tree of 217bbd7 with only this change applied.
  • Three mutants each fail the new test:
    • a count cap of 8 (captures 0 12 0 4 4 against 0 12 0 0 0);
    • nodes not counted (0 held);
    • nodes not capped (16 graphs, 128 nodes over a cap of 100).
  • Real decode graphs on the 5080: qwen3-0.6b is 538 nodes, and Ternary Bonsai 2 27B's two are 1,192 and 1,096.

…, not their count

A CUDA graph's executable holds device memory for each node of its graph, from a pool of the driver's that never
shrinks. On an RTX 5080 (driver 610.43), executables of 1 to 2,048 kernel nodes took 1.9 to 3.0 KiB a node, at 16 B
to 2 KiB of kernel parameters. Destroying them gave nothing back to cudaMemGetInfo, and instantiating them again took
nothing more. Under a cap on the nodes held, executables of 20 to 4,000 nodes were instantiated and evicted 3,500 times
over, and the pool stayed where the most held at once had put it: 14 MiB at 4,096 nodes, 44 at 16,384, 88 at 32,768.

d0f8bae capped their count at 8. A -sm tensor split computes one graph per all-reduce step on each device, two a
layer every token, and each is its own CUDA graph. Under a cap of 8 each was evicted before its turn came round again,
then warmed up and captured anew every token. qwen3-0.6b Q4_0 (llama-bench tg128, -ts 1/1, RTX 5080 + 5070 Ti)
decoded at 257.0 tok/s with -sm tensor against 558.0 uncapped.

GGML_CUDA_GRAPH_NODES (default 32768, 0 for no cap) now bounds the nodes a context's executables hold: the least
recently used are evicted before a new one is instantiated. GGML_CUDA_GRAPH_MAX still bounds their count, but only when
set (default 0, no cap). The nodes of a captured graph are counted (cudaGraphGetNodes). The log gives the nodes held
beside the count, and the device memory all the executables took; it logs at INFO when that memory grows or the nodes
held reach a power of two, and at DEBUG otherwise. A -sm tensor context holds hundreds of graphs, each a new most.

Decode graphs on the 5080: qwen3-0.6b's is 538 nodes, and Ternary Bonsai 2 27B's two are 1,192 and 1,096 (8 MiB in
all), so the default holds 27 of the 27B's.

test-cuda-graph-cap-nodes sets a cap of 100 nodes, with graphs of eight nodes (x halved eight times) that are built
once and computed again.
- Twelve computed in turn, five times round: each is captured once, then replayed from the third time round (captures
  0 12 0 0 0). All twelve are held, 96 nodes.
- Sixteen in turn: at most 96 nodes are held.
It fails under a cap of 8 graphs (captures 0 12 0 4 4, 8 held), when the nodes are not counted (0 held), and when they
are not capped (16 graphs, 128 nodes). test-cuda-graph-cap still passes under a cap of 4 graphs, and so do
test-cuda-graph-key and test-cuda-graph-src-type.
@marcospaulo
marcospaulo merged commit d38768c into main Sep 28, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant