Summary
On a native Windows build with the HIP backend (gfx1151, Strix Halo), a single
/v1/chat/completions request completes prefill and then kills the server with
ACCESS_VIOLATION (0xc0000005) before the first decode token. Deterministic.
The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_async dereferences AsyncOutputSlot::device_sampled_ids
on the CPU. The async path is entered because RocmBackend advertises
SupportsAsyncSampledTokenReadback() == true unconditionally, although the contract that
predicate stands for (a HIP mirror or a D2H copy of dev_ids) is not implemented.
Environment
- Windows 11 (build 26200), Strix Halo (Radeon 8060S),
gfx1151, iGPU / UMA, 128 GB shared
- vllm.cpp commit
636926736c0b6053eda99d08c1f753b29938fac0 (also reproduced on
9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6)
- built natively for Windows: HIP + hipBLASLt via official TheRock 10.0
(C:\TheRock\10.0.0-official\build), clang/lld from TheRock, -O3 -DNDEBUG -g -Xclang -gcodeview, vllm.cpp 0.0.3 c-abi=29
- model:
Qwen3.5-4B-Q4_K_M GGUF (bartowski), --max-model-len 2048, --max-num-seqs 1
Steps to reproduce
vllm-server.exe --model <path>\Qwen3.5-4B-Q4_K_M.gguf \
--host 127.0.0.1 --port 18125 --served-model-name smoke \
--device auto --max-model-len 2048 --gpu-memory-utilization 0.35 \
--max-num-seqs 1 --max-num-batched-tokens 512 \
--disable-metrics --no-enable-thinking --verbose
then one request:
POST /v1/chat/completions
{"model":"smoke","messages":[{"role":"user","content":"Reply with the single word OK"}],"max_tokens":16}
Expected
A chat completion with generated tokens.
Actual
Prefill completes (18/18 tokens), then the process dies. Connection closed, no response.
Windows Application Error: 0xc0000005, faulting process vllm-server.exe
(observed twice on 9e63db5d at the same offset, once on 6369267).
Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
(edd8.10d18): Access violation - code c0000005
READ_ADDRESS: 00000006b08cb000
ExceptionAddress: vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
FAILURE_BUCKET: NULL_POINTER_READ_c0000005_vllm-server.exe!vllm::v1::GPUModelRunner::sample_tokens_async
!vprot 0x6b08cb000
State: 00002000 MEM_RESERVE
Protect: 00000001 PAGE_NOACCESS
Type: 00020000 MEM_PRIVATE
RegionSize: 000000045b620000 (17.428 GB)
Faulting instruction (rdi = 0, rax = 0x6b08cb000):
140fb6180: movq 0x4f0(%rbp), %rax ; rax = dev_ids (device allocation)
140fb6187: movl (%rax,%rdi,8), %eax ; <-- AV: int64 read from device memory on the CPU
140fb6194: movl %eax, (%rcx,%rdi,4) ; last_sampled_tokens[i] = (int32)ids[i]
Symbolicated stack:
vllm_server!vllm::v1::GPUModelRunner::sample_tokens_async+0x687
vllm_server!vllm::v1::EngineCore::step_with_batch_queue
vllm_server!vllm::v1::EngineCoreProc::process_engine_step
vllm_server!vllm::v1::EngineCoreProc::run_busy_loop
vllm_server!vllm::v1::InprocClient (inlined)
vllm_server!std::thread::_Invoke<...>
kernel32!BaseThreadInitThunk
ntdll!RtlUserThreadStart
Root cause
The device-resident decode path:
src/vllm/v1/worker/gpu/runner.cpp:5569-5573 — the slot's device buffer is handed to the
sampler, which writes the argmax ids device-resident.
- Both device write-back branches are inside
#ifdef VLLM_CPP_CUDA
(runner.cpp:5604-5661), so a HIP build compiles them out.
- The fallback
else branch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"
and does:
vt::GetBackend(dev.type).Synchronize(queue_);
const int64_t* ids = static_cast<const int64_t*>(dev_ids); // 5669: device pointer
for (int i = 0; i < num_reqs; ++i) {
...
input_batch_.last_sampled_tokens[i] = static_cast<int32_t>(ids[i]); // 5680-5681
}
dev_ids is AsyncOutputSlot::device_sampled_ids, allocated by vt::Alloc — on ROCm that
is hipMalloc memory, which the CPU may not dereference on Windows (WDDM gives the process
a GPU VA that is MEM_RESERVE/PAGE_NOACCESS), hence the AV.
That branch is reachable because the capability gate passes:
src/vllm/v1/worker/gpu/runner.cpp:112-115 — QueueSupportsAsyncInputCombine() asks
backend->SupportsAsyncSampledTokenReadback(); runner.cpp:518-519 turns that into
async_input_combine_.
src/vt/rocm/rocm_backend.hip:328:
bool SupportsAsyncSampledTokenReadback() const override { return true; }
while include/vt/backend.h:217-220 documents the precondition:
TODO(rocm): an INTEGRATED non-CUDA GPU reports UnifiedMemory()==true … such a backend
may override this true once a HIP sampled-token mirror or a D2H copy of dev_ids lands.
- the same backend answers, in the same file:
bool UnifiedMemory() const override { return unified_memory_; } // 589
bool DeviceMemoryIsHostAddressable() const override { return unified_memory_; } // 606
and on this part unified_memory_ is false, because the memory policy withholds the
managed branch (hipDeviceAttributePageableMemoryAccess = 0) — see
include/vt/rocm/rocm_arch.h:156-162 and rocm_backend.hip:219-221.
So RocmBackend simultaneously claims "the host may not dereference Alloc pointers" and
"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.
Why Linux ROCm does not see this
Where the managed allocator branch is taken (hipMallocManaged, PageableMemoryAccess = 1,
UnifiedMemory() == true) the device pointer is host-addressable, so the same host read
silently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.
Workaround (verified)
VT_ASYNC_RUNNER=0 on the same binary, model and port:
HTTP_OK 'OK' | prefill 18/18 | 2 tokens | finish_reason=stop
The crash reproduces with the default (async on) 3/3 across the two commits.
Suggested fix
- Minimal/honest capability:
SupportsAsyncSampledTokenReadback() should not be a constant
true on ROCm — at least return unified_memory_; (same source as
DeviceMemoryIsHostAddressable()), so the async combine is not engaged where no
mirror/D2H exists.
- Implement the contract for HIP: reuse
DownloadCommittedIds (runner.cpp:5234, already
used at runner.cpp:5720 — copy queue + fork/ready events + staging buffer) or read
slot->pinned_host after ready_event, exactly as
AsyncGPUModelRunnerOutput::get_output() does
(src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencing dev_ids
on the host.
- Defensive: gate the host read at
runner.cpp:5669 on
backend.DeviceMemoryIsHostAddressable() and fail with a named check instead of a raw AV.
Notes on the Windows build recipe (context, not part of the bug)
Getting a native Windows HIP link at all required: force-linking vllm.lib with
-Xlinker /WHOLEARCHIVE (self-registering platform TUs are otherwise dropped, which shows
up earlier as fatal: vt: no platform registered for device type 0), removing a
Strawberry/MinGW -lpthreads contamination from PATH, and forcing the Windows thread
cache values. With those, --target server builds cleanly; a full all-target build still
fails in tools/bench/conv1d_scaling_probe.cpp (#include <sys/resource.h>).
Artifacts
- full WER minidump available on request (5.5 GB, not attachable): stack above is from it
vllm-server.pdb (155,807,744 B) exists for the 6369267 build, so the dump symbolizes
directly
Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).
Summary
On a native Windows build with the HIP backend (
gfx1151, Strix Halo), a single/v1/chat/completionsrequest completes prefill and then kills the server withACCESS_VIOLATION (0xc0000005)before the first decode token. Deterministic.The fault is a host read of a device allocation: the non-CUDA write-back in
GPUModelRunner::sample_tokens_asyncdereferencesAsyncOutputSlot::device_sampled_idson the CPU. The async path is entered because
RocmBackendadvertisesSupportsAsyncSampledTokenReadback() == trueunconditionally, although the contract thatpredicate stands for (a HIP mirror or a D2H copy of
dev_ids) is not implemented.Environment
gfx1151, iGPU / UMA, 128 GB shared636926736c0b6053eda99d08c1f753b29938fac0(also reproduced on9e63db5dd33b35e7cc57d0f0e80fe6c7d5ababa6)(
C:\TheRock\10.0.0-official\build), clang/lld from TheRock,-O3 -DNDEBUG -g -Xclang -gcodeview,vllm.cpp 0.0.3 c-abi=29Qwen3.5-4B-Q4_K_MGGUF (bartowski),--max-model-len 2048,--max-num-seqs 1Steps to reproduce
then one request:
Expected
A chat completion with generated tokens.
Actual
Prefill completes (
18/18tokens), then the process dies. Connection closed, no response.Windows Application Error:
0xc0000005, faulting processvllm-server.exe(observed twice on
9e63db5dat the same offset, once on6369267).Crash evidence (WER LocalDump + cdb, symbols from the build's own PDB)
Faulting instruction (
rdi = 0,rax = 0x6b08cb000):Symbolicated stack:
Root cause
The device-resident decode path:
src/vllm/v1/worker/gpu/runner.cpp:5569-5573— the slot's device buffer is handed to thesampler, which writes the argmax ids device-resident.
#ifdef VLLM_CPP_CUDA(
runner.cpp:5604-5661), so a HIP build compiles them out.elsebranch (runner.cpp:5662-5684) is labelled "HOST path (CPU backend)"and does:
dev_idsisAsyncOutputSlot::device_sampled_ids, allocated byvt::Alloc— on ROCm thatis
hipMallocmemory, which the CPU may not dereference on Windows (WDDM gives the processa GPU VA that is
MEM_RESERVE/PAGE_NOACCESS), hence the AV.That branch is reachable because the capability gate passes:
src/vllm/v1/worker/gpu/runner.cpp:112-115—QueueSupportsAsyncInputCombine()asksbackend->SupportsAsyncSampledTokenReadback();runner.cpp:518-519turns that intoasync_input_combine_.src/vt/rocm/rocm_backend.hip:328:while
include/vt/backend.h:217-220documents the precondition:and on this part
unified_memory_is false, because the memory policy withholds themanaged branch (
hipDeviceAttributePageableMemoryAccess = 0) — seeinclude/vt/rocm/rocm_arch.h:156-162androcm_backend.hip:219-221.So
RocmBackendsimultaneously claims "the host may not dereferenceAllocpointers" and"async sampled-token readback is supported", and the runner relies on the latter to
host-dereference a device buffer.
Why Linux ROCm does not see this
Where the managed allocator branch is taken (
hipMallocManaged,PageableMemoryAccess = 1,UnifiedMemory() == true) the device pointer is host-addressable, so the same host readsilently succeeds and the missing D2H is masked. On Windows/gfx1151 the branch is withheld
(issue #2511 policy), so the read faults.
Workaround (verified)
VT_ASYNC_RUNNER=0on the same binary, model and port:The crash reproduces with the default (async on) 3/3 across the two commits.
Suggested fix
SupportsAsyncSampledTokenReadback()should not be a constanttrueon ROCm — at leastreturn unified_memory_;(same source asDeviceMemoryIsHostAddressable()), so the async combine is not engaged where nomirror/D2H exists.
DownloadCommittedIds(runner.cpp:5234, alreadyused at
runner.cpp:5720— copy queue + fork/ready events + staging buffer) or readslot->pinned_hostafterready_event, exactly asAsyncGPUModelRunnerOutput::get_output()does(
src/vllm/v1/worker/gpu/async_output.cpp:113-126), instead of dereferencingdev_idson the host.
runner.cpp:5669onbackend.DeviceMemoryIsHostAddressable()and fail with a named check instead of a raw AV.Notes on the Windows build recipe (context, not part of the bug)
Getting a native Windows HIP link at all required: force-linking
vllm.libwith-Xlinker /WHOLEARCHIVE(self-registering platform TUs are otherwise dropped, which showsup earlier as
fatal: vt: no platform registered for device type 0), removing aStrawberry/MinGW
-lpthreadscontamination fromPATH, and forcing the Windows threadcache values. With those,
--target serverbuilds cleanly; a full all-target build stillfails in
tools/bench/conv1d_scaling_probe.cpp(#include <sys/resource.h>).Artifacts
vllm-server.pdb(155,807,744 B) exists for the6369267build, so the dump symbolizesdirectly
Related: #1627 (the TT backend never advertises this same capability — the opposite
direction of the same predicate).