Skip to content

feat(BACKEND-CUDA-SM120): opt-in per-allocation trace - #3358

Closed
iiLaurens wants to merge 2 commits into
mudler:mainfrom
iiLaurens:row/BACKEND-CUDA-SM120
Closed

iiLaurens wants to merge 2 commits into
mudler:mainfrom
iiLaurens:row/BACKEND-CUDA-SM120

Conversation

@iiLaurens

@iiLaurens iiLaurens commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Row

BACKEND-CUDA-SM120 — one row per PR. Spec:
.agents/specs/cuda-alloc-trace.md.

Before starting

  • Issue/PR search: no issue existed for per-allocation visibility;
    VT_CUDA_ALLOC_STATS is aggregate by design and stays as it is. Filed
    ISSUE-LOCAL-01M3QFZD0PTVP2CT9S2RYG8M2H.
  • Pull request shape selected at row claim: one PR for the spec and the
    implementation.
  • Anchors inspected: src/vt/cuda/cuda_backend.cu (Check at :55-59,
    Alloc/Free at :120-140, StatsEnabled, DeviceMemoryInfo).

What changed

VT_CUDA_ALLOC_TRACE=1 prints every cudaMalloc of at least 16 MiB with its
size, running live bytes and free device memory; every free of the same size; and
the request that fails, with its error — before Check throws. Below 2 GiB free
every allocation prints regardless of size, so the endgame is never filtered out.
Zero-cost when unset: one process-static env read and one size comparison on a
path that is about to call the driver.

The instrument exists for the row's card: a 24 GiB sm_120a serving a ~20 GiB
NVFP4 arm out of the same pool as the KV cache and the repack scratch, where the
question is the allocation ORDER and the live curve, not the process total.

Evidence

$ ./build-cuda/tests/test_cuda_alloc_trace
test cases:  1 |  1 passed | 0 failed | 0 skipped
assertions: 16 | 16 passed | 0 failed

Build: -DVLLM_CPP_CUDA=ON -DVLLM_CPP_CUDA_ARCHITECTURES=120a, local sm_120a
(RTX PRO 4000 Blackwell, 24 GiB).

REAL RUN, on the artifact the instrument exists for — the official ModelOpt NVFP4
checkpoint ukisai/Swift-1.5-Qwen3.8-27b-NVFP4 (21.94 GB) on the local
sm_120a 24 GiB card, which currently dies in the NVFP4 Marlin repack without
PR #3355's leak fix:

$ VT_CUDA_ALLOC_TRACE=1 ./build-cuda/examples/vllm-cli --model <dir> --device cuda \
    --kv-cache-dtype fp8 --max-num-seqs 1 --kv-cache-memory 1000000000 \
    --max-tokens 4 --prompt "The capital of France is" --temperature 0
[cuda-free]  size=2425.0 MiB live=6.74 GiB free=13.81 GiB
engine-fatal: EngineCore busy loop threw: vt cuda: marlin_repack: malloc bqt: out of memory

716 trace lines on the run, ending in the live/free curve above. The trace turned a
one-line OOM into the state at the failure: 6.74 GiB live, 13.81 GiB reported
free, the last frees being 2.4 GiB and 606 MiB blocks.

And it named the instrument's own blind spot, which the log makes visible: the
failing call is cudaMallocAsync in src/vt/cuda/cuda_marlin_repack.cu:161, a
kernel-local stream-ordered pool allocation that does NOT go through
CudaBackend::Alloc, so no [cuda-alloc] FAILED line is printed for it. That is
recorded under "Honest gaps" and in the spec's ## Owed, not papered over.

MUTATION, the guarantee this pins — force AllocTraceEnabled() false and
rebuild (the kernel still compiles):

$ mutation: AllocTraceEnabled() returns false
ERROR: CHECK( alloc_line.find("[cuda-alloc]") != std::string::npos )
ERROR: CHECK( alloc_line.find("size=32.0 MiB") != std::string::npos )
ERROR: CHECK( free_line.find("[cuda-free]") != std::string::npos )
ERROR: CHECK( failed_line.find("FAILED") != std::string::npos )
test cases: 1 | 0 passed | 1 failed | 0 skipped
assertions: 16 | 9 passed | 7 failed

The sentinel assertion keeps the capture honest: with the redirect broken the
case fails on the sentinel rather than reporting "the trace printed nothing".

The case allocates 32 MiB, frees it, and requests an impossible size; it captures
stderr around each and asserts the [cuda-alloc] / [cuda-free] / FAILED
lines carry the right size. It asserts its own capture with a sentinel, so a
redirect that caught nothing fails rather than reading as "the trace printed
nothing".

  • commit style + trailer contracts OK; Preflight on this box is red at BASE (35 gates: record drift, the
    e126687a9a vs a7c23ac96d oracle pin mismatch, check-test-registration's
    CMake probe, and the missing file tool); each failing gate re-run
    on a clean origin/main worktree fails there too.

  • tests that cover this change: tests/vt/test_cuda_alloc_trace.cpp (new)

  • docs/ENVIRONMENT.md documents VT_CUDA_ALLOC_TRACE in this change

Speed claims

  • This PR makes NO speed claim. It is a diagnostic: it prints, and it is
    compiled out of the hot path when unset (the guard is a process-static
    bool read plus a size compare).

Honest gaps

  • The flag-off arm cannot be exercised in the same binary: the flag is
    process-static by design, and the test binary enables it before main. The
    off path is the same guard with AllocTraceEnabled() false; every CUDA suite
    in the tree runs it.
  • The instrument traces CudaBackend::Alloc/Free. Allocations that bypass the
    backend are invisible: on a real run the failing call was cudaMallocAsync in
    src/vt/cuda/cuda_marlin_repack.cu:161 (the Marlin repack's bqt scratch), so
    the log shows the state at the failure but not the failing request's own line.
    Extending the trace to the pool APIs, or routing those sites through the
    backend, is the named next step (spec ## Owed).
  • A second real limitation the run exposed: the repack's cudaMallocAsync failed
    with 13.81 GiB reported free, which points at stream-ordered-pool growth rather
    than total exhaustion. On the same card the same command succeeds once
    PR fix(LOAD-MODELOPT-NVFP4-BORROW): one owner per device resident #3355 frees the leaked fp4 originals, so the leak is the practical cause;
    the pool interaction itself is not characterized here.
    Preflight on this box is red at BASE (35 gates: record drift, the
    e126687a9a vs a7c23ac96d oracle pin mismatch, check-test-registration's
    CMake probe, and the missing file tool); each failing gate re-run
    on a clean origin/main worktree fails there too.

Refs ISSUE-LOCAL-01M3QFZD0PTVP2CT9S2RYG8M2H

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:opencode-go/deepseek-v4.1-flash [pi]

dev added 2 commits September 29, 2026 21:09
A first-forward OOM on a 24 GiB sm_120a card names one failing size, and nothing
in the tree can say which allocation sequence produced the pressure. Record the
instrument, its threshold rationale, the capture-sentinel test, and what it
does not prove.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:opencode-go/deepseek-v4.1-flash [pi]
VT_CUDA_ALLOC_TRACE=1 prints each cudaMalloc of at least 16 MiB with its
size, live bytes and free device memory, each free of the same size, and
the request that fails. A first-forward OOM becomes a size list instead
of a guess. Zero-cost when unset.

The threshold is the signal-to-noise line: the pool hands out MiB-scale
blocks, and below 2 GiB free every allocation prints regardless of size so
the endgame is never missed. The trace is keyed by pointer, so a free of
an untracked block is ignored.

The new test in test_cuda_alloc_trace captures stderr around a 32 MiB
allocation, its free, and a request that cannot fit. It asserts its own
capture with a sentinel, so a redirect that caught nothing fails instead
of reading as "the trace printed nothing". The flag is process-static, so
the binary enables it before main and is flag-ON by construction; the
flag-off path is the same guard and is covered by every CUDA suite, none
of which sets the variable.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:opencode-go/deepseek-v4.1-flash [pi]
@iiLaurens
iiLaurens marked this pull request as ready for review September 29, 2026 21:37

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two source-level regressions need fixes at d416d6a:

  1. Windows build: tests/vt/test_cuda_alloc_trace.cpp:32 calls ::setenv unconditionally in the global initializer. The target is registered unconditionally, so a CPU-only MSVC build must compile that POSIX call before the runtime CUDA skip can help. Please use the existing tests/support/test_env.h helper, which supplies the _putenv_s arm. Also guard the alloc_line assertions on capture support: the non-POSIX branch leaves that string empty and still asserts that it contains trace output.
  2. Combined diagnostics: CudaBackend::Alloc returns from the trace branch before Stats().mallocs is incremented, but Free still increments Stats().frees. With both VT_CUDA_ALLOC_TRACE=1 and VT_CUDA_ALLOC_STATS=1, one allocation/free pair is therefore counted as zero allocations and one free. This contradicts the spec's promise that aggregate statistics remain unchanged. Keep allocation accounting on the common success path and add a regression with both flags enabled.

I read the full diff, the allocation/free/statistics paths, the test-target helper, and the existing portable environment helper. These are static findings, not a claim of executed Windows/CUDA tests: this host lacks Python, CMake, and a C++ compiler. @mudler @richiejp: please hold the merge recommendation until these paths have coverage.

@iiLaurens

Copy link
Copy Markdown
Contributor Author

Withdrawing this change: VT_CUDA_ALLOC_TRACE is a debugging convenience, not a
product capability, and it is not worth the review budget. The findings here (the
aggregate-statistics accounting and the unconditional setenv) are specific to
this change and are moot with it, so there is nothing to correct.

No further commits are pushed; the branch is left as the reviewed head.

@iiLaurens iiLaurens closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants