perf(cuda): cap the CUDA graphs a context keeps by their nodes, not their count - #77
Merged
Merged
Conversation
…, not their count A CUDA graph's executable holds device memory for each node of its graph, from a pool of the driver's that never shrinks. On an RTX 5080 (driver 610.43), executables of 1 to 2,048 kernel nodes took 1.9 to 3.0 KiB a node, at 16 B to 2 KiB of kernel parameters. Destroying them gave nothing back to cudaMemGetInfo, and instantiating them again took nothing more. Under a cap on the nodes held, executables of 20 to 4,000 nodes were instantiated and evicted 3,500 times over, and the pool stayed where the most held at once had put it: 14 MiB at 4,096 nodes, 44 at 16,384, 88 at 32,768. d0f8bae capped their count at 8. A -sm tensor split computes one graph per all-reduce step on each device, two a layer every token, and each is its own CUDA graph. Under a cap of 8 each was evicted before its turn came round again, then warmed up and captured anew every token. qwen3-0.6b Q4_0 (llama-bench tg128, -ts 1/1, RTX 5080 + 5070 Ti) decoded at 257.0 tok/s with -sm tensor against 558.0 uncapped. GGML_CUDA_GRAPH_NODES (default 32768, 0 for no cap) now bounds the nodes a context's executables hold: the least recently used are evicted before a new one is instantiated. GGML_CUDA_GRAPH_MAX still bounds their count, but only when set (default 0, no cap). The nodes of a captured graph are counted (cudaGraphGetNodes). The log gives the nodes held beside the count, and the device memory all the executables took; it logs at INFO when that memory grows or the nodes held reach a power of two, and at DEBUG otherwise. A -sm tensor context holds hundreds of graphs, each a new most. Decode graphs on the 5080: qwen3-0.6b's is 538 nodes, and Ternary Bonsai 2 27B's two are 1,192 and 1,096 (8 MiB in all), so the default holds 27 of the 27B's. test-cuda-graph-cap-nodes sets a cap of 100 nodes, with graphs of eight nodes (x halved eight times) that are built once and computed again. - Twelve computed in turn, five times round: each is captured once, then replayed from the third time round (captures 0 12 0 0 0). All twelve are held, 96 nodes. - Sixteen in turn: at most 96 nodes are held. It fails under a cap of 8 graphs (captures 0 12 0 4 4, 8 held), when the nodes are not counted (0 held), and when they are not capped (16 graphs, 128 nodes). test-cuda-graph-cap still passes under a cap of 4 graphs, and so do test-cuda-graph-key and test-cuda-graph-src-type.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The shared branch's commits after #76:
-sm tensor. The meta backend computes a graph per all-reduce step on each device, two a layer. Each was evicted before its reuse and captured anew every token: qwen3-0.6b decoded at 257.0 tok/s against 558.0 uncapped on a 5080 + 5070 Ti.GGML_CUDA_GRAPH_NODES(default 32768, about 88 MiB) now bounds the nodes held, evicting the least recently used.GGML_CUDA_GRAPH_MAXbounds the count only when set.GGML_CUDA_GRAPH_NODES=0 GGML_CUDA_GRAPH_MAX=8restores d0f8bae.Checks:
test-cuda-graph-cap-nodes(new) andtest-cuda-graph-cap,test-cuda-graph-key,test-cuda-graph-src-typepass. They ran in a clean tree of 217bbd7 with only this change applied.