Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 27 additions & 0 deletions .agents/issues/KV-FP8/ISSUE-LOCAL-01M3QG4WWC3X0C84M9PWQZAJ10.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
ID: ISSUE-LOCAL-01M3QG4WWC3X0C84M9PWQZAJ10
Title: fp8 KV prefill runs the scalar CUDA-core flash kernel at 3.7x the bf16 cache, because FA-2 admits only bf16 q/KV/out while the fp8 store presents f32
Row: KV-FP8
State: OPEN
Kind: feature
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-29
Updated: 2026-09-29
Closed: -

## Problem

MEASURED on the 27B NVFP4 arm at 8k prefill: 28.0s with the fp8 KV cache against 7.8s with the same kernel on a bf16 cache, a 3.7x gap. nsys attributes it entirely to the kernel choice: 25.1s of 36s sits in `PagedFlashKernel<float, unsigned char, ...>`, the scalar CUDA-core arm, while the bf16 store runs the FA-2 split-KV kernel at 2.9 ms/layer. CAUSE: the model presents f32 q/out for a non-bf16 store (`GdnOutDType`/the KV-store route), and the vendored FA-2 admission requires bf16 q, bf16 KV and bf16 out, so every non-bf16 cache falls off the tensor-core ladder onto the per-element fp8 dequant. The W2 CUDA arm of this row (issue #1593) landed the fp8 store and the read dequant for CORRECTNESS, so this is the prefill-PERFORMANCE half the row never had. THE FIX: dequantize the fp8 cache ONCE per layer into a dense bf16 scratch and run the normal bf16 dispatch on it. The scratch is one block per request with block_size = max_seq and an identity block table, so the kernel s paged address IS the dense address and NO attention kernel changes. A shared `KvCachePresentsBf16` helper presents bf16 for an fp8 store, so FA-2 admits with zero cast kernels and any model that gains an fp8 store inherits the CUDA path unchanged. `VT_ATTN_FP8_DENSE=0` restores the per-read dequant for a same-binary A/B. TWO EARLIER SHAPES OF THE SAME LEVER WERE TRIED AND REJECTED, and the record keeps them: (a) routing fp8 through the bf16 WMMA ladder with an fp8->bf16 cast inside the K/V staging (8k 28.0 -> 9.1s) still pays a per-element dequant on every re-stream, and (b) converting the staging with `__nv_cvt_fp8_to_halfraw` made that conversion cheap (bit-identical) but did not remove it. The dense dequant removes it entirely for prefill. Numbers after: 32k prefill 34.9s / 938 tok/s against 36.4s / 901 for the bf16 store and 39.9s / 822 for llama.cpp s q8_0 KV, at 20.5 GiB against 22.5; 8k 7.6s / 1000 tok/s; decode with MTP n=3 38.6 vs 37.4. Op-level parity against the f32 reference is 1.9e-6 max abs err (the exact dequantized values), all 33 paged-attention cases pass, and benchmarks/paged_attn_prefill_ab.cpp is the isolated sweep that separated the kernel cost from the engine.

## Resolution

- 2026-09-30: review repairs on the W1 PR (mudler/vllm.cpp#3360). The dense bf16
scratch and the identity table are owned by stream-aware scope guards, so the
steps that can throw between the two allocations (identity copy, dequant
launch, the bf16 dispatch, the vectors between them) release both buffers on
the same stream. A new case in `test_ops_paged_attn` injects the identity
allocation failure and the identity copy failure and checks the CUDA pool's
used bytes return to the pre-call value without the guard masking the
original exception. The benchmark header no longer claims a bf16 parity
comparison it does not perform; parity is gated by the test suite.
96 changes: 96 additions & 0 deletions .agents/specs/fp8-kv-prefill-dense-dequant.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
# fp8 KV prefill on the bf16 dispatch: one dense dequant per layer — ISSUE-LOCAL-01M3QG4WWC3X0C84M9PWQZAJ10

The fp8 KV cache read sent every prefill to the scalar CUDA-core flash kernel,
3.7x slower than the same kernel on a bf16 cache, because FA-2 admits only bf16
q/KV/out and the fp8 store presents f32. This is the prefill-performance half of
`KV-FP8`; W2 (issue [#1593](https://github.com/mudler/vllm.cpp/issues/1593))
landed the fp8 store and read-dequant for correctness.

Issue: [ISSUE-LOCAL-01M3QG4WWC3X0C84M9PWQZAJ10](../issues/KV-FP8/ISSUE-LOCAL-01M3QG4WWC3X0C84M9PWQZAJ10.md).
Owning row: `KV-FP8` ([engine-matrix.md](../engine-matrix.md)); the row's spec is
[fp8-kv-cache.md](fp8-kv-cache.md).

## Premise, grounded

| Where (line anchors at this branch's base, `b45a94273`) | What |
|---|---|
| `src/vt/cuda/cuda_paged_attn.cu:148-197` | The fp8 K/V read with the per-element dequant folded in (`Fp8E4M3ToF32Dev`), the W2 arm. |
| `src/vt/cuda/cuda_paged_attn.cu:2909`, `:2954` | The dispatch comments that name the split: the tensor-core ladder for bf16, the f32-q/out scalar arm otherwise. |
| `include/vllm/model_executor/models/kv_cache_route.h:40,73-78` | The store/read route that hands `kv_cache_dtype` to the backend. |
| `src/vllm/model_executor/models/qwen3_5.cpp` (the attention preamble) | The model-side arm this change routes. |
| `tests/vt/test_ops_paged_attn.cpp` | 33 cases on the base, including the W2 fp8 parity case. |

Measured on the 27B NVFP4 arm at 8k prefill: 28.0 s fp8 vs 7.8 s bf16 (3.7x),
25.1 s of 36 s in `PagedFlashKernel<float, unsigned char, …>`, against
2.9 ms/layer on the FA-2 split-KV kernel for a bf16 store.

## Design

Dequantize the fp8 cache ONCE per layer into a dense bf16 scratch, then run the
shipped bf16 dispatch on it:

- The scratch is **one block per request with `block_size = max_seq` and an
identity block table**, so the kernel's paged address IS the dense address and
no attention kernel changes.
- Each scratch allocation is owned by a scope guard from the moment it
succeeds. The identity allocation and copy, the dequant launch, the bf16
dispatch it feeds, and the vectors between them can all throw, and the guard
frees on the SAME stream, so the free is ordered behind the work that reads
the buffer. The success path still checks its own frees explicitly.
- The new shared `KvCachePresentsBf16` helper presents bf16 for an fp8 store, so
FA-2 admits with zero cast kernels. It is model-agnostic: any model that gains
an fp8 store inherits the CUDA path unchanged.
- `VT_ATTN_FP8_DENSE=0` restores the per-read dequant for a same-binary A/B.

**Two rejected levers are recorded rather than deleted**, because both are
plausible and neither survives the numbers:

1. Routing fp8 through the bf16 WMMA ladder with an fp8→bf16 cast inside the K/V
staging (8k 28.0 → 9.1 s, `VT_ATTN_FP8_WMMA`). It keeps the per-element
dequant on every re-stream, so the gap to bf16 (7.8 s) only narrows.
2. Converting the staging with `__nv_cvt_fp8_to_halfraw` instead of the software
decode — bit-identical and cheaper, but still per re-stream.

The dense dequant removes the per-read dequant entirely for prefill, which is
why it is the shipped shape. The bench that separated kernel cost from engine
cost is `benchmarks/paged_attn_prefill_ab.cpp`, landed with this change.

## Tests

`tests/vt/test_ops_paged_attn.cpp` gains the fp8-dense parity case: the fp8 cache
against the f32 reference at **1.9e-6 max abs err** — the exact dequantized
values, tighter than the bf16-compute envelope — with the rest of the 33-case
suite unchanged. A second case injects the two failures a healthy device cannot
produce on demand (the identity-table allocation, and the identity-table copy
after both allocations succeeded) and asserts that the CUDA memory pool's used
bytes return to their pre-call value after the exception unwinds. It also
asserts the propagated message is the failing `Check`, so the guard's destructor
cannot mask the original exception.

## What this does NOT claim

- **No speed claim in this PR.** The numbers above are the author's measurement
on the local `sm_120a` card with host embedding on (`VT_HOST_EMBEDDING=1`), and
they are recorded as the shape's evidence, not as an operator gate. A
same-binary A/B under the GPU lease is owed by the operator, as the helper
template requires.
- The 20.5 GiB-vs-22.5 GiB peak comparison and the llama.cpp q8_0 denominator
(39.9 s / 822 tok/s) are the author's run of the same recipe; they are not
re-measured here.

## Gates

- `ctest --test-dir build -R test_ops_paged_attn` on a CUDA build
(`-DVLLM_CPP_CUDA=ON -DVLLM_CPP_CUTLASS_FETCH=ON` for FA-2).
- `VT_ATTN_FP8_DENSE=0` returns the base behaviour (the A/B arm).
- The full `ctest --test-dir build` on the same build.
- `scripts/agent-preflight.sh --staged`.

## Owed

- The operator's same-binary A/B under lease, with the 27B NVFP4 arm, both arms
in one binary, and the recipe written into `docs/BENCHMARKS.md`.
- A `sm_120a` re-measurement: the author's numbers are from the local consumer
card, and the fleet gate model runs elsewhere.
- Execution of the scratch-ownership case, which needs a CUDA device; the CUDA
lane is its gate.
8 changes: 8 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -2905,6 +2905,14 @@ target_include_directories(vllm_music3_vocoder_conv_ab SYSTEM PRIVATE
target_compile_features(vllm_music3_vocoder_conv_ab PRIVATE cxx_std_20)
vllm_cpp_set_warnings(vllm_music3_vocoder_conv_ab)

# Paged-attention prefill A/B (paged vs dense KV source, bf16 vs fp8 cache).
# The isolated sweep behind the fp8 KV prefill dense dequant (KV-FP8): one
# vt::PagedAttention call per arm, so the number is the kernel cost alone.
add_executable(vllm_paged_attn_prefill_ab benchmarks/paged_attn_prefill_ab.cpp)
target_link_libraries(vllm_paged_attn_prefill_ab PRIVATE vllm)
target_compile_features(vllm_paged_attn_prefill_ab PRIVATE cxx_std_20)
vllm_cpp_set_warnings(vllm_paged_attn_prefill_ab)

# ── The `vt::Conv1d` decomposition probe (#672, #1334) ───────────────────────
# `tools/bench/conv1d_scaling_probe.cpp` is the instrument behind the scaling
# curve and the residency ablation in `.agents/specs/vt-conv1d-time-block.md`.
Expand Down
188 changes: 188 additions & 0 deletions benchmarks/paged_attn_prefill_ab.cpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,188 @@
// Paged-attention PREFILL A/B: paged vs dense KV source, bf16 vs fp8 cache, at
// the 27B full-attention shape (hq=32, hk=4, d=256, block_size=32).
//
// WHY AN ISOLATED SWEEP EXISTS: the engine-level fp8 prefill numbers (28s vs 8s
// at 8k) conflate the attention kernel, the per-layer dequant scratch, the D2H
// sync and the driver allocator. This runs ONE vt::PagedAttention call per arm
// on synthetic data, so the number is the kernel cost and nothing else.
//
// ARMS (one process per arm; the dispatch knobs are process-static):
// BENCH_KV=bf16 paged bf16 cache (the bf16 engine baseline)
// BENCH_KV=fp8 paged fp8 cache (dense scratch ON by default)
// BENCH_KV=fp8 VT_ATTN_FP8_DENSE=0 per-read dequant inside the kernel
// BENCH_KV=bf16 BENCH_DENSE=1 dense bf16 + identity table (layout control)
//
// ONE arm per process, so this binary cannot compare arms to each other. It
// prints a checksum of the arm's output, which makes a wildly wrong arm visible
// in the logs; it is not a parity gate. Parity is gated by
// `tests/vt/test_ops_paged_attn.cpp`, where each arm is compared against an f32
// reference on the exact dequantized values (< 5e-2 max abs err).
#include <algorithm>
#include <chrono>
#include <cmath>
#include <cstdint>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <string>
#include <vector>

#include "vt/backend.h"
#include "vt/dtype.h"
#include "vt/fp8_kv.h"
#include "vt/ops.h"

using vt::Backend;
using vt::DeviceType;
using vt::DType;
using vt::Fp8KVCacheDataType;
using vt::PagedAttentionArgs;
using vt::Queue;
using vt::Tensor;

namespace {

Tensor MakeT(void* data, DType dt, const std::vector<int64_t>& shape) {
Tensor t;
t.data = data;
t.dtype = dt;
t.device = vt::Device{DeviceType::kCUDA, 0};
t.rank = static_cast<int>(shape.size());
int64_t stride = 1;
for (int i = t.rank - 1; i >= 0; --i) {
t.shape[i] = shape[static_cast<size_t>(i)];
t.stride[i] = stride;
stride *= shape[static_cast<size_t>(i)];
}
return t;
}

std::vector<float> RandF32(size_t n, uint32_t seed) {
std::vector<float> v(n);
uint32_t s = seed;
for (auto& x : v) {
s = s * 1664525u + 1013904223u;
x = (static_cast<float>(s >> 8) / static_cast<float>(1u << 24)) * 4.0f - 2.0f;
}
return v;
}

struct Buf {
Backend& b;
void* p = nullptr;
size_t bytes = 0;
Buf(Backend& backend, size_t n) : b(backend), bytes(n) { p = b.Alloc(n == 0 ? 1 : n); }
~Buf() { b.Free(p); }
Buf(const Buf&) = delete;
Buf& operator=(const Buf&) = delete;
};

} // namespace

int main(int argc, char** argv) {
const int64_t T = argc > 1 ? std::atoll(argv[1]) : 8192;
const int64_t Hq = argc > 2 ? std::atoll(argv[2]) : 32;
const int64_t Hk = argc > 3 ? std::atoll(argv[3]) : 4;
const int64_t D = argc > 4 ? std::atoll(argv[4]) : 256;
const int64_t BS = argc > 5 ? std::atoll(argv[5]) : 32;
const int reps = argc > 6 ? std::atoi(argv[6]) : 3;
const char* kv_env = std::getenv("BENCH_KV");
const bool fp8 = kv_env != nullptr && std::strcmp(kv_env, "fp8") == 0;
const char* dense_env = std::getenv("BENCH_DENSE");
const bool dense_bf16 = !fp8 && dense_env != nullptr && dense_env[0] == '1';
const float scale = std::pow(static_cast<float>(D), -0.5f);
const int64_t N = T / BS; // one request, identity block order

Backend& gpu = vt::GetBackend(DeviceType::kCUDA);
Queue q = gpu.CreateQueue();

auto q_host = RandF32(static_cast<size_t>(T * Hq * D), 2024);
auto k_host = RandF32(static_cast<size_t>(N * BS * Hk * D), 137);
auto v_host = RandF32(static_cast<size_t>(N * BS * Hk * D), 179);
std::vector<int32_t> block_table(static_cast<size_t>(N));
for (int64_t i = 0; i < N; ++i) block_table[static_cast<size_t>(i)] = static_cast<int32_t>(i);
std::vector<int32_t> seq_lens = {static_cast<int32_t>(T)};
std::vector<int32_t> qsl = {0, static_cast<int32_t>(T)};

Buf dq(gpu, static_cast<size_t>(T * Hq * D) * sizeof(float));
Buf dbt(gpu, block_table.size() * sizeof(int32_t));
Buf dsl(gpu, seq_lens.size() * sizeof(int32_t));
Buf dqsl(gpu, qsl.size() * sizeof(int32_t));
Buf dout(gpu, static_cast<size_t>(T * Hq * D) * sizeof(float));
gpu.Copy(q, dq.p, q_host.data(), dq.bytes);
gpu.Copy(q, dbt.p, block_table.data(), dbt.bytes);
gpu.Copy(q, dsl.p, seq_lens.data(), dsl.bytes);
gpu.Copy(q, dqsl.p, qsl.data(), dqsl.bytes);

// Cache: paged [N, BS, Hk, D] (bf16 or fp8), or dense [1, T, Hk, D] bf16.
const int64_t cache_elems = dense_bf16 ? T * Hk * D : N * BS * Hk * D;
Buf kcache(gpu, static_cast<size_t>(cache_elems) * (fp8 ? 1 : 2));
Buf vcache(gpu, static_cast<size_t>(cache_elems) * (fp8 ? 1 : 2));
Tensor kt, vt;
if (fp8) {
std::vector<uint8_t> k8(k_host.size()), v8(v_host.size());
for (size_t i = 0; i < k_host.size(); ++i) {
k8[i] = vt::StoreKvFp8E4M3(k_host[i], 1.0f);
v8[i] = vt::StoreKvFp8E4M3(v_host[i], 1.0f);
}
gpu.Copy(q, kcache.p, k8.data(), k8.size());
gpu.Copy(q, vcache.p, v8.data(), v8.size());
kt = MakeT(kcache.p, DType::kI8, {N, BS, Hk, D});
vt = MakeT(vcache.p, DType::kI8, {N, BS, Hk, D});
} else {
std::vector<uint16_t> kb(k_host.size()), vb(v_host.size());
for (size_t i = 0; i < k_host.size(); ++i) {
kb[i] = vt::F32ToBF16(k_host[i]);
vb[i] = vt::F32ToBF16(v_host[i]);
}
gpu.Copy(q, kcache.p, kb.data(), kb.size() * 2);
gpu.Copy(q, vcache.p, vb.data(), vb.size() * 2);
const std::vector<int64_t> shape =
dense_bf16 ? std::vector<int64_t>{1, T, Hk, D} : std::vector<int64_t>{N, BS, Hk, D};
kt = MakeT(kcache.p, DType::kBF16, shape);
vt = MakeT(vcache.p, DType::kBF16, shape);
}

// Dense arm: one block per request, block_size = T, identity table.
int32_t* bt_ptr = static_cast<int32_t*>(dbt.p);
const int64_t bt_elems = N;
const int64_t block_size = dense_bf16 ? T : BS;

PagedAttentionArgs args{scale, /*causal=*/true};
if (fp8) {
args.kv_cache_dtype = Fp8KVCacheDataType::kFp8E4M3;
args.k_scale = 1.0f;
args.v_scale = 1.0f;
}

Tensor q_t = MakeT(dq.p, DType::kF32, {T, Hq, D});
Tensor out_t = MakeT(dout.p, DType::kF32, {T, Hq, D});
Tensor bt_t = MakeT(bt_ptr, DType::kI32, {1, dense_bf16 ? 1 : bt_elems});
Tensor sl_t = MakeT(dsl.p, DType::kI32, {1});
Tensor qsl_t = MakeT(dqsl.p, DType::kI32, {2});

std::vector<float> got(static_cast<size_t>(T * Hq * D));
double ms = 0.0;
for (int r = 0; r < reps + 2; ++r) {
const auto t0 = std::chrono::steady_clock::now();
vt::PagedAttention(q, out_t, q_t, kt, vt, bt_t, sl_t, qsl_t, args);
gpu.Synchronize(q);
const double dt =
std::chrono::duration<double, std::milli>(std::chrono::steady_clock::now() - t0).count();
if (r >= 2) ms += dt;
}
ms /= reps;
gpu.Copy(q, got.data(), dout.p, got.size() * sizeof(float));
gpu.Synchronize(q);

const char* arm = fp8 ? (std::getenv("VT_ATTN_FP8_DENSE") != nullptr ? "fp8/per-read" : "fp8/dense")
: (dense_bf16 ? "bf16/dense" : "bf16/paged"); std::printf("arm=%-12s T=%lld hq=%lld hk=%lld d=%lld bs=%lld %.1f ms %.0f tok/s\n", arm,
static_cast<long long>(T), static_cast<long long>(Hq), static_cast<long long>(Hk),
static_cast<long long>(D), static_cast<long long>(block_size), ms, T / (ms / 1000.0));
// Checksum so a wrong-but-fast arm is visible.
double sum = 0.0;
for (float x : got) sum += x;
std::printf(" checksum=%.6f\n", sum);
gpu.DestroyQueue(q);
return 0;
}
9 changes: 9 additions & 0 deletions include/vllm/model_executor/models/kv_cache_route.h
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,15 @@ inline bool IsFp8KvCache(const PagedKvCache& kv) {
return fp8_kind;
}

// The dtype the ATTENTION presents for this store. An fp8 store is read through
// the bf16 dense dequant scratch (`Fp8PagedToDenseBf16Kernel`), so it presents
// bf16 exactly as a bf16 store does; a model that decides its query/out dtype
// from the store's dtype must ask THIS, or an fp8 store silently falls off the
// bf16 attention lanes (FA-2 included) onto the per-read CUDA-core path.
inline bool KvCachePresentsBf16(const PagedKvCache& kv) {
return kv.dtype == vt::DType::kBF16 || IsFp8KvCache(kv);
}

// The KV STORE. `k`/`v` are the model-dtype [T, Hkv, Dh] tensors the attention
// preamble produced; `k_cache`/`v_cache` are this layer's `KvSlice` views.
//
Expand Down
7 changes: 5 additions & 2 deletions src/vllm/model_executor/models/qwen3_5.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -6033,7 +6033,10 @@ DBuf FullAttnBlockPaged(Dev d, const FullAttnLayerWeights& w, const HfConfig& cf
/*num_reqs=*/meta.num_reqs,
/*uniform_spec_query_len=*/meta.uniform_spec_query_len,
/*causal=*/meta.causal,
/*kv_cache_bf16=*/kv.dtype == DType::kBF16,
// An fp8 store is served through a bf16 dense scratch (the prefill
// dequant), so it presents bf16 too — that is what admits the FA-2
// prefill lane. The decode lanes keep their own admission.
/*kv_cache_bf16=*/dense_attn::KvCachePresentsBf16(kv),
/*kv_block_multiple_16=*/kv.block_size % 16 == 0,
/*preamble_with_cos_sin=*/FuseAttnPreambleOn(fp4) && sdi.has_attn_cos_sin,
/*fa2_platform=*/fa2_platform,
Expand Down Expand Up @@ -6114,7 +6117,7 @@ DBuf FullAttnBlockPaged(Dev d, const FullAttnLayerWeights& w, const HfConfig& cf
Tensor vw = v3;
DBuf kbf(d, DType::kBF16, {T, Hkv, Dh});
DBuf vbf(d, DType::kBF16, {T, Hkv, Dh});
if (kv.dtype == DType::kBF16 || dense_attn::IsFp8KvCache(kv)) {
if (dense_attn::KvCachePresentsBf16(kv)) {
// K may already be bf16 (an FA2 preamble emits bf16 k directly —
// the RN round of the same f32 value this CastBf16 would produce); only
// down-cast when the preamble/fallback produced f32 K.
Expand Down
Loading