From d463a1a41c98306646e2e5d6f72aec899df7c3cd Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=E5=88=98=E5=AE=89?= Date: Wed, 23 Sep 2026 12:54:19 +0800 Subject: [PATCH] cuda: raise CUDA graph cache cap from 64 to 512 for tensor parallel MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The meta backend (-sm tensor) dispatches one subgraph per AllReduce split point to every simple backend. For qwen4exp that is ~128 subgraphs per device, each getting its own entry in the per-device cuda_graphs map. With max_cuda_graphs = 64 the LRU cap evicted entries on every decode token, so graph->uid never matched and warmup never completed: zero CUDA graph captures, fully eager decode (~14.6k kernel launches per token per device). Measured on 8x V100-SXM2 (NVLink), qwen4exp Q4_K_XL: TP4 tg64: 6.77 -> 44.41 ± 6.26 t/s (tg512: 47.93) TP8 tg64: 5.50 -> 30.37 ± 5.86 t/s nsys tg64: cudaGraphLaunch 0 -> 48888, cudaLaunchKernel 2.96M -> 118k Layer-split mode only needs one entry per device and was unaffected. 512 covers large TP graphs; the existing 10s idle sweep still bounds the map size in steady state. Co-Authored-By: Claude Code --- ggml/src/ggml-cuda/common.cuh | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/ggml/src/ggml-cuda/common.cuh b/ggml/src/ggml-cuda/common.cuh index 068c9b15f10b..9ac72a7194a2 100644 --- a/ggml/src/ggml-cuda/common.cuh +++ b/ggml/src/ggml-cuda/common.cuh @@ -1467,7 +1467,13 @@ struct ggml_backend_cuda_context { #ifdef USE_CUDA_GRAPH std::unordered_map> cuda_graphs; - static const size_t max_cuda_graphs = 64; + // NOTE: tensor-parallel (-sm tensor, meta backend) models dispatch one subgraph per AllReduce + // split point to every simple backend (~2 per layer for qwen4exp -> >100 keys per device). + // The old cap of 64 made the LRU evict entries every token, permanently resetting CUDA graph + // warmup and forcing fully eager decode (V100 x8, qwen4exp Q4_K_XL: TP4 6.8 -> 44 t/s after + // this change; nsys: cudaGraphLaunch 0 -> 48888 over a tg64 run). 512 covers large TP graphs; + // the 10s idle sweep below still bounds memory for steady state. + static const size_t max_cuda_graphs = 512; int64_t last_graph_eviction_sweep = 0;