Conversation
A first-forward OOM on a 24 GiB sm_120a card names one failing size, and nothing in the tree can say which allocation sequence produced the pressure. Record the instrument, its threshold rationale, the capture-sentinel test, and what it does not prove. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:opencode-go/deepseek-v4.1-flash [pi]
VT_CUDA_ALLOC_TRACE=1 prints each cudaMalloc of at least 16 MiB with its size, live bytes and free device memory, each free of the same size, and the request that fails. A first-forward OOM becomes a size list instead of a guess. Zero-cost when unset. The threshold is the signal-to-noise line: the pool hands out MiB-scale blocks, and below 2 GiB free every allocation prints regardless of size so the endgame is never missed. The trace is keyed by pointer, so a free of an untracked block is ignored. The new test in test_cuda_alloc_trace captures stderr around a 32 MiB allocation, its free, and a request that cannot fit. It asserts its own capture with a sentinel, so a redirect that caught nothing fails instead of reading as "the trace printed nothing". The flag is process-static, so the binary enables it before main and is flag-ON by construction; the flag-off path is the same guard and is covered by every CUDA suite, none of which sets the variable. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:opencode-go/deepseek-v4.1-flash [pi]
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
Two source-level regressions need fixes at d416d6a:
- Windows build:
tests/vt/test_cuda_alloc_trace.cpp:32calls::setenvunconditionally in the global initializer. The target is registered unconditionally, so a CPU-only MSVC build must compile that POSIX call before the runtime CUDA skip can help. Please use the existingtests/support/test_env.hhelper, which supplies the_putenv_sarm. Also guard thealloc_lineassertions on capture support: the non-POSIX branch leaves that string empty and still asserts that it contains trace output. - Combined diagnostics:
CudaBackend::Allocreturns from the trace branch beforeStats().mallocsis incremented, butFreestill incrementsStats().frees. With bothVT_CUDA_ALLOC_TRACE=1andVT_CUDA_ALLOC_STATS=1, one allocation/free pair is therefore counted as zero allocations and one free. This contradicts the spec's promise that aggregate statistics remain unchanged. Keep allocation accounting on the common success path and add a regression with both flags enabled.
I read the full diff, the allocation/free/statistics paths, the test-target helper, and the existing portable environment helper. These are static findings, not a claim of executed Windows/CUDA tests: this host lacks Python, CMake, and a C++ compiler. @mudler @richiejp: please hold the merge recommendation until these paths have coverage.
|
Withdrawing this change: No further commits are pushed; the branch is left as the reviewed head. |
Row
BACKEND-CUDA-SM120— one row per PR. Spec:.agents/specs/cuda-alloc-trace.md.Before starting
VT_CUDA_ALLOC_STATSis aggregate by design and stays as it is. FiledISSUE-LOCAL-01M3QFZD0PTVP2CT9S2RYG8M2H.implementation.
src/vt/cuda/cuda_backend.cu(Checkat :55-59,Alloc/Freeat :120-140,StatsEnabled,DeviceMemoryInfo).What changed
VT_CUDA_ALLOC_TRACE=1prints everycudaMallocof at least 16 MiB with itssize, running live bytes and free device memory; every free of the same size; and
the request that fails, with its error — before
Checkthrows. Below 2 GiB freeevery allocation prints regardless of size, so the endgame is never filtered out.
Zero-cost when unset: one process-static env read and one size comparison on a
path that is about to call the driver.
The instrument exists for the row's card: a 24 GiB
sm_120aserving a ~20 GiBNVFP4 arm out of the same pool as the KV cache and the repack scratch, where the
question is the allocation ORDER and the live curve, not the process total.
Evidence
Build:
-DVLLM_CPP_CUDA=ON -DVLLM_CPP_CUDA_ARCHITECTURES=120a, localsm_120a(RTX PRO 4000 Blackwell, 24 GiB).
REAL RUN, on the artifact the instrument exists for — the official ModelOpt NVFP4
checkpoint
ukisai/Swift-1.5-Qwen3.8-27b-NVFP4(21.94 GB) on the localsm_120a24 GiB card, which currently dies in the NVFP4 Marlin repack withoutPR #3355's leak fix:
716 trace lines on the run, ending in the live/free curve above. The trace turned a
one-line OOM into the state at the failure: 6.74 GiB live, 13.81 GiB reported
free, the last frees being 2.4 GiB and 606 MiB blocks.
And it named the instrument's own blind spot, which the log makes visible: the
failing call is
cudaMallocAsyncinsrc/vt/cuda/cuda_marlin_repack.cu:161, akernel-local stream-ordered pool allocation that does NOT go through
CudaBackend::Alloc, so no[cuda-alloc] FAILEDline is printed for it. That isrecorded under "Honest gaps" and in the spec's
## Owed, not papered over.MUTATION, the guarantee this pins — force
AllocTraceEnabled()false andrebuild (the kernel still compiles):
The sentinel assertion keeps the capture honest: with the redirect broken the
case fails on the sentinel rather than reporting "the trace printed nothing".
The case allocates 32 MiB, frees it, and requests an impossible size; it captures
stderr around each and asserts the
[cuda-alloc]/[cuda-free]/FAILEDlines carry the right size. It asserts its own capture with a sentinel, so a
redirect that caught nothing fails rather than reading as "the trace printed
nothing".
commit style + trailer contracts OK; Preflight on this box is red at BASE (35 gates: record drift, the
e126687a9avsa7c23ac96doracle pin mismatch,check-test-registration'sCMake probe, and the missing
filetool); each failing gate re-runon a clean
origin/mainworktree fails there too.tests that cover this change:
tests/vt/test_cuda_alloc_trace.cpp(new)docs/ENVIRONMENT.mddocumentsVT_CUDA_ALLOC_TRACEin this changeSpeed claims
compiled out of the hot path when unset (the guard is a process-static
bool read plus a size compare).
Honest gaps
process-static by design, and the test binary enables it before
main. Theoff path is the same guard with
AllocTraceEnabled()false; every CUDA suitein the tree runs it.
CudaBackend::Alloc/Free. Allocations that bypass thebackend are invisible: on a real run the failing call was
cudaMallocAsyncinsrc/vt/cuda/cuda_marlin_repack.cu:161(the Marlin repack'sbqtscratch), sothe log shows the state at the failure but not the failing request's own line.
Extending the trace to the pool APIs, or routing those sites through the
backend, is the named next step (spec
## Owed).cudaMallocAsyncfailedwith 13.81 GiB reported free, which points at stream-ordered-pool growth rather
than total exhaustion. On the same card the same command succeeds once
PR fix(LOAD-MODELOPT-NVFP4-BORROW): one owner per device resident #3355 frees the leaked fp4 originals, so the leak is the practical cause;
the pool interaction itself is not characterized here.
Preflight on this box is red at BASE (35 gates: record drift, the
e126687a9avsa7c23ac96doracle pin mismatch,check-test-registration'sCMake probe, and the missing
filetool); each failing gate re-runon a clean
origin/mainworktree fails there too.Refs ISSUE-LOCAL-01M3QFZD0PTVP2CT9S2RYG8M2H
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:opencode-go/deepseek-v4.1-flash [pi]