Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
ID: ISSUE-LOCAL-01M3V01R90QQTQ08BWSBM7AGN2
Title: kCausalConv1dFwd (GDN prefill conv) has no native Vulkan kernel
Row: BACKEND-VULKAN
State: CLOSED
Kind: gap
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-01
Updated: 2026-10-01
Closed: 2026-10-01

## Problem

vt::CausalConv1dFwd, the GDN PREFILL depthwise causal conv, is not registered for DeviceType::kVULKAN, so every GDN prefill on a Vulkan queue runs it on the portable CPU reference tier, once per GDN layer, behind a full batch drain. src/vt/vulkan/vulkan_ops.cpp left it as a follow-up because the op reads the OLD conv-state window and overwrites it in the same call; the comment names the two safe dispatch shapes (a serial invocation per (sequence, channel) over the whole token range, or a buffered old row). .agents/specs/vulkan-full-support.md section 6.0a lists it as one of the two remaining reference-tier declines.

## Resolution

Closed by row/BACKEND-VULKAN-CONVFWD: native vt_causal_conv1d_fwd, default one invocation per (sequence, channel) serial over its tokens with the old window copied before the write-back; opt-in token split (VT_VULKAN_CONV_TARGET_GROUPS) giving every t < width to block 0; unservable shapes decline through GetOpFallback; VT_VULKAN_CONV_FWD=0 keeps the reference tier. Gated by tests/vt/test_vulkan_backend.cpp against the CPU oracle; the split arm is red on llvmpipe with the pre-fix plain split and green with this PR.
48 changes: 48 additions & 0 deletions .agents/specs/vulkan-full-support.md
Original file line number Diff line number Diff line change
Expand Up @@ -613,6 +613,54 @@ gate is a memcmp of the two arms in ONE process, plus the specialization VALUES
from `PipelineKeys()`, because both arms are the same module and produce
identical bytes; a numeric check alone could never see the mechanism.

### 6.0b `VK-G` partial: the PREFILL CONV landed — 2026-10-01

`row/BACKEND-VULKAN-CONVFWD`. `kCausalConv1dFwd` is native
(`vt_causal_conv1d_fwd`); module count 43 -> 44. This updates the closing
sentence of §6.0a: of the two ops it named as the reference-tier declines on that
path, `kRopeCosSinCache` remains, by design.

**Why it had been left.** The op computes every output from the OLD conv-state
window and then overwrites that window in the same call, so a dispatch that
splits a sequence across invocations reads state another invocation is
rewriting. `vulkan_ops.cpp` named the two safe shapes: a serial invocation per
(sequence, channel) over the whole token range, or a buffered old row.

**What landed.** The first of those shapes is the DEFAULT: one invocation per
(sequence, channel), the CPU kernel's own `ForRows(n * c_dim, ...)` unit, with
the old window copied into a private array before the write-back (the CPU
kernel's `old_row`). Per-element arithmetic is ported 1:1 from
`src/vt/cpu/cpu_ops.cpp` `CausalConv1dFwdKernel`, including its silu spelling.

An OPT-IN token split, `VT_VULKAN_CONV_TARGET_GROUPS`, divides each sequence
into token blocks in the same dispatch. Only `t < width` reads the carried state;
the shader gives every such token to block 0, which also writes the state back
and is the only block that touches `conv_state`. Without that rule the split
raced: on llvmpipe the wrong elements were exactly the `t < width` tokens that
fell outside block 0. It is off by default; whether a split is worth anything is
a property of the device.

Shapes the shader does not serve DECLINE to the reference tier through
`GetOpFallback`, the seam §6.0 uses: K = 1, a kernel width past its 8-slot
window, `conv_state` rows wider than K-1 (the CPU reference addresses rows with
stride K-1 while the op layer admits wider rows; declining keeps the reference's
answer instead of choosing), storage dtypes outside f32/f16/bf16, a
`has_initial_state` that is neither i8 nor i32, a grid past the device's
`maxComputeWorkGroupCount[0]`, and any index that would not fit the shader's
uint32 arithmetic. The native path does not read `query_start_loc` on the host;
the shader's guard keeps every access in bounds but does not validate the table,
so a malformed table that the CPU reference would reject gives unspecified
output here.

**Gates.** `test_vulkan_backend`, against the CPU oracle: the default mapping
and the split on a varlen batch (lengths 5, 1, 9, one sequence without initial
state), with the specialization value asserted; i8 flags at byte offsets 1-3, no
bias, silu off, bf16 and f16 operands with bf16 or f32 output, a zero-length
sequence and a padded `x` row stride; a bf16
`conv_state` against the f32 arm; and the K = 10 and widened-row declines.
Outputs to the GDN NMSE tolerance, rolled state bit-exact.
`VT_VULKAN_CONV_FWD=0` keeps the op on the reference tier for a same-binary A/B.

### 6.0a `VK-G` partial: the FUSED ATTN PREAMBLE landed — 2026-08-09

`row/BACKEND-VULKAN-QKNORM`. `kAttnQkNormRopeGate` — gemma-RMSNorm(q) +
Expand Down
2 changes: 2 additions & 0 deletions docs/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -292,6 +292,8 @@ portable/reference path. In normal operation leave them unset.
| `VT_VULKAN_RMSNORM` | auto | Which `vt::RmsNorm` SPIR-V module runs: `wide` forces the 1024-invocation subgroup-reducing one, `base` forces the portable 128-invocation one, unset lets the device capability decide (1024 invocations on the X axis plus compute subgroup BASIC and ARITHMETIC). The wide module exists because `RmsNorm` dispatches ONE WORKGROUP PER ROW and a batch-1 decode step has exactly one row: on Qwen3.6-27B that put 128 invocations on a 5120-wide row, four warps of one SM, with the rest of the GPU idle. MEASURED on GB10 by the two-length GPU-timestamp diff: **0.0611 -> 0.0123 ms/call, 7.88 -> 1.59 ms/token**, and paired decode **241.9 -> 235.6 ms** median TPOT. The tell that it was OCCUPANCY and not the reduction is that the SAME shader costs 0.066 ms/call during PREFILL, where 32 rows give it 32 workgroups and 32x the data. It exists for the same-binary A/B and so the unit gate can exercise the fallback on hardware that would always pick the wide arm. Vulkan-only |
| `VT_VULKAN_MATMUL_NCOLS` | 4 | Output columns each lane of the portable scalar GEMM computes, in the `[K,N]` (non-transposed) orientation only. At 1 the kernel is the flat one-invocation-per-output-element body; above 1 a workgroup takes `128*NCOLS` CONSECUTIVE output columns of one row, so at each step of K it reads a contiguous run of that many elements instead of 128. This is the ONE decode GEMM that cannot reach the `vt_matmul_vec` tactic, because in `[K,N]` the lanes are already coalesced and the GEMV shape would make them strided; on the 27B it is the lm_head, `m=1 k=5120 n=248320`, 2.54 GB moved per token. MEASURED on GB10, 27B decode, `ms/call` medians over interleaved replicates: NCOLS 1 = 12.48, 2 = 12.46, **4 = 11.54**, 8 = 12.81, with 4 winning **6 of 6** interleaved pairs against 1. Blocking is a TRADE, not a monotone win: at 8 the dispatch falls to 243 workgroups (~31k threads) and the device runs out of work to hide memory latency with faster than the longer contiguous run buys back. It rides a specialization constant, so every arm is the same committed module and they A/B in one binary. Every arm is BIT-IDENTICAL -- each accumulator owns one output element and sums the whole K sequentially, which is the CPU kernel's order -- so this kernel keeps the byte-exact tier that the coopmat and GEMV tactics gave up; a memcmp gates that. Vulkan-only |
| `VT_VULKAN_COOPMAT` | on | `=0` forces the Vulkan GEMM onto the portable SCALAR kernel instead of the cooperative-matrix (tensor-core) tactic. The coopmat path is selected only where the device reports the exact `16x16x16 bf16/bf16/f32/f32 SUBGROUP` configuration, subgroup size is 32, both operands are bf16, and M, N and K are all multiples of 16. The whole-tile requirement on M and N is not a tuning choice: `coopMatLoad` reads a full 16x16 tile with no masking, so a partial tile reads past the operand and can fault the GPU. Ragged shapes fall back to the scalar kernel; this switch bypasses that selection entirely. It exists for the same-binary A/B in `examples/vulkan-gemm-ab` (measured 11.1x-32.9x on NVIDIA Thor) and as the bisect lever if a coopmat result is ever suspect. Vulkan-only |
| `VT_VULKAN_CONV_FWD` | on | `=0` leaves `vt::CausalConv1dFwd` (the GDN PREFILL conv) on the portable CPU reference tier instead of registering the native `vt_causal_conv1d_fwd` kernel, which is where the op ran before that kernel existed. Read once, at backend registration. It exists for the same-binary A/B and as the bisect lever, the same shape as `VT_VULKAN_COOPMAT`. The native kernel matches the CPU reference to a tolerance on the outputs and bit-exactly on the rolled `conv_state`. Vulkan-only |
| `VT_VULKAN_CONV_TARGET_GROUPS` | unset (no split) | Opt-in TOKEN SPLIT for the native prefill conv. Unset, each (sequence, channel) is one invocation serial over its tokens. Set to a workgroup target `g` (1..65536), each sequence is split into `ceil(g / base_groups)` token blocks, where `base_groups` is the default grid in workgroups, capped at the mean sequence length, at what one dispatch can launch, and at 64 (each block count is its own compiled pipeline). Only tokens `t < width` read the carried state; the shader gives all of them to block 0, which is also the only block that writes the state back, so no other block touches `conv_state`. Each element is computed by the same expression whatever the block count; `test_vulkan_backend` compares 1 and 5 blocks bit-for-bit. Whether a split is worth anything depends on the device, which is why there is no default. Read on every call. Vulkan-only |
| `VT_GLM5_NEXT_DEVICE_EXPERTS` | **off (opt-in)** | `=1` lets `Glm5NextForConditionalGeneration` (GLM-5.3-Flash) accept a non-CPU queue and route its routed-expert keep-quant GEMM to the device, against banks made resident by `dense_attn::ResidentWeight`. **IT IS OFF BECAUSE THE PATH IT ENABLES IS MEASURED TO CRASH, not because it is unmeasured.** On `dgx:gpu0` against the published 101.24 GiB `UD-Q2_K_XL` artifact, ALL THREE `--device cuda` legs died with SIGSEGV (rc=139) having emitted no token, interleaved against three `--device cpu` legs that all emitted ` Paris.` from the same binary. **The mixed-residency reading of those legs is FALSIFIED and the cause is elsewhere**: the two log lines that suggested it are once-flags, and the process dies in `StoreCaches`, which host-stores into the runner's `cudaMalloc` KV pages after the forward has already returned. That defect is older than this knob and only became reachable when the non-CPU refusal above it was removed; see `.agents/specs/glm5-next-flash.md` O49 and [#2480](https://github.com/mudler/vllm.cpp/issues/2480), which owns the fix. The default is the refusal the tree carried before the arm existed, because turning a clean named error into a segfault is strictly worse for a user. **Set this only to debug that crash; it is not a serving knob.** Parsed strictly (`1` and nothing else, not the usual first-character rule) precisely because it opts into a crashing path. Inert on every other model and on `--device cpu`. See `.agents/specs/glm5-next-flash.md` O46 and [#2464](https://github.com/mudler/vllm.cpp/issues/2464) |
| `VT_GLM5_NEXT_DEVICE` | **off (opt-in)** | `=1` routes the entire GLM-5.3-Flash forward through `Glm5NextDeviceForward`, which dispatches embedding, RMSNorm, MoE combine, the k-pool indexer and `lm_head` through `vt::*` device ops on the queue, keeping MLA attention, mHC sites and the dense MLP as host-fallback islands (the `kimi_linear_device.cpp` single-queue pattern). This is a superset of `VT_GLM5_NEXT_DEVICE_EXPERTS`: when on, the whole forward delegates and the per-arm experts split is not reached. On a CPU queue the `vt::*` kernels use float32 accumulation where the host reference uses double, so the output agrees within a float-vs-double envelope rather than byte-exact. Inert on every other model. See `.agents/specs/glm5-next-flash.md` W9c-3 and [#3175](https://github.com/mudler/vllm.cpp/pull/3175) |

Expand Down
Loading