Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-ENG-HOST-EMBEDDING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-ENG-HOST-EMBEDDING

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-ENG-HOST-EMBEDDING` | `ENG-HOST-EMBEDDING` (`ACTIVE`) | pi (deepseek-v4.1-flash), helper role — holds the row's spec, the host arm and its test; a fresh reviewer over the immutable head still owes its own pass | isolated worktree `/tmp/wt-embed`, Release CPU build plus the local `sm_120a` card (RTX PRO 4000 Blackwell, 24 GiB, a NON-fleet device). NO fleet lease, NO oracle run, NO benchmark, NO model weights loaded | `row/ENG-HOST-EMBEDDING`, PR [#3356](https://github.com/mudler/vllm.cpp/pull/3356) | Owns ONLY: `.agents/specs/host-embedding.md`, the `ENG-HOST-EMBEDDING` row and its two counts in `.agents/engine-matrix.md`, this claim, issue `ISSUE-LOCAL-01M3QETSTM8X8AJKBM9QCTGBKM`, and the implementation it carries (`host_embedding.{h,cpp}`, the `EmbedGather` call sites, `docs/ENVIRONMENT.md`, `tests/vllm/models/test_host_embedding.cpp`). EXCLUDES every other row's records and every model/capability the seam is merely wired into; EXCLUDES the end-to-end VRAM/throughput measurement, which the row owes | `ACTIVE` | 2026-09-29 — row claimed; spec, implementation, test and docs land in ONE pull request per the recorded Git-integration preference |
5 changes: 3 additions & 2 deletions .agents/engine-matrix.md

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
ID: ISSUE-LOCAL-01M3QETSTM8X8AJKBM9QCTGBKM
Title: The token table has no host-resident arm: every dense forward uploads [vocab, H] to the device even though the gather could run in host RAM
Row: ENG-HOST-EMBEDDING
State: OPEN
Kind: feature
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-29
Updated: 2026-09-29
Closed: -

## Problem

A dense forward that owns a device embedding table pays [vocab, H] of device memory for it and gathers there. On the cards this project targets that is real: a 151k x 5120 bf16 table is 1.5 GiB of a 24 GiB pool, and on a tied-head model the table is kept for the lm_head GEMM regardless. The gather itself is tiny on the host: one row per id through the CPU vt::Embedding kernel, then one [T, H] copy to the device, which is llama.cpp`s ggml_get_rows shape (ggml/src/ggml-cpu/ops.cpp:4850 @ b10451) rather than vLLM`s device-side gather. Nothing in the tree offered this: `ResidentWeight`/`EmbedGather` always uploaded, so a text-only or memory-bound serve had no way to keep the table host-side. The change adds `VT_HOST_EMBEDDING=1` (host gather, one row per id, bf16/f16/f32 and every GGUF block format, bf16 byte-copy fast path) behind the existing `EmbedGather` seam, so every dense forward that owns a device table gets it for free and the device behaviour is unchanged when the flag is off or the host bytes are gone. It is vllm.cpp-original, so the secondary oracle is llama.cpp`s ggml_get_rows, and the correctness bar is the same pinned IQ4_NL/Q5_0 golden vectors the device-side vt::Embedding gate already uses.

## Resolution

- 2026-09-30: review repairs on the W1 PR (mudler/vllm.cpp#3356). The host arm's
async-id override now routes through the shared `detail::ApplyDeviceTokenIds`
body instead of its own unchecked Copy, so an override longer than the embed
input is refused rather than written past the `[T]` buffer; two cases in
`tests/vllm/models/test_host_embedding.cpp` pin the oversized refusal and the
shorter-prefix tail preservation. The test's global initializer uses the
portable `vllm_test::SetEnv` (`tests/support/test_env.h`) instead of POSIX
`::setenv`, which an unconditional target cannot compile under MSVC.
149 changes: 149 additions & 0 deletions .agents/specs/host-embedding.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
# Host-resident token table: gather the embedding rows on the CPU — ISSUE-LOCAL-01M3QETSTM8X8AJKBM9QCTGBKM

A dense forward that owns a device embedding table uploads `[vocab, H]` once and
gathers there. On a 24 GiB card that is 1.5 GiB of pool for a 151k x 5120 bf16
table, and on a tied head the table is kept for its GEMM whatever the gather does.
The gather itself is small: one row per id.

Issue: [ISSUE-LOCAL-01M3QETSTM8X8AJKBM9QCTGBKM](../issues/ENG-HOST-EMBEDDING/ISSUE-LOCAL-01M3QETSTM8X8AJKBM9QCTGBKM.md).
Owning row: `ENG-HOST-EMBEDDING` ([engine-matrix.md](../engine-matrix.md)), added
with this spec.

## Scope

IN: the host-resident arm of the token-table gather — the `VT_HOST_EMBEDDING`
switch, the CPU gather through `vt::Embedding`, the `[T,H]` copy to the device,
the `EmbedGather` seam that owns both arms, the wake-up of every dense forward
that owns a device table, `docs/ENVIRONMENT.md`, and the tests that pin the arm.

OUT: the tied-head GEMM and every other consumer of the table (the table's host
bytes stay available for them); the decode-graph arms whose ids live on a device
tensor (they keep the device gather by design); any change to the device arm's
bytes or dtype; and the end-to-end VRAM/throughput measurement, which is a row
deliverable recorded under `## Owed`, not a precondition for the arm.

## Upstream chain

**vLLM has no equivalent.** Its embedding is a device-resident
`VocabParallelEmbedding` parameter and the gather runs on the device. This is a
vllm.cpp-original capability, so the primary oracle has nothing to mirror and the
shape comes from a **secondary oracle**: llama.cpp's `ggml_get_rows` /
`ggml_compute_forward_get_rows_q` (`ggml/src/ggml-cpu/ops.cpp:4850` @ `b10451`),
the dequantizing host gather that reads ONE ROW per id and never materializes the
table. That is the same function `tests/vt/test_ops_embedding_quant.cpp` cites for
the device-side `vt::Embedding` op, so both arms are compared against one
denominator.

## Our baseline

Before this change every dense forward that owns a device table staged it with
`ResidentWeight(d, table, {vocab, H})` and gathered with `vt::Embedding(d.q, ...)`.
For an fp8/non-CUDA-device path `ResidentWeight` aliases the host bytes, but on a
staging backend (CUDA, non-unified) it uploads `[vocab, H]` and keeps it resident
for the process. There was no way to keep the table host-side, and no seam that
owned both arms.

## Port map

| Piece | Where |
|---|---|
| The seam and the flag | `include/vllm/model_executor/models/host_embedding.h`, `src/vllm/model_executor/models/host_embedding.cpp` |
| The wake-up call sites | the Qwen3.5 family, MuseGlimmer, the shared Qwen3 dense driver, and the classic dense families (Gemma 1-4, GLM4, Granite, MiniCPM 1/3, OLMo2, OPT, Phi, Phi3, StableLM, Command-R, DeepSeek-V2, GLM-MoE-DSA, Dots3-Note, Nemotron-H, Voxtral) |
| The async id override consumed before the gather | `src/vllm/model_executor/models/qwen3_5_internal.h` (`detail::TakeDeviceTokenIds`, then `detail::ApplyDeviceTokenIds` for the splice) |
| The documented knobs | `docs/ENVIRONMENT.md` (`VT_HOST_EMBEDDING`, `VT_HOST_EMBED_TRACE`) |
| Build registration | the new TU in `CMakeLists.txt` |

## Tests to port

The oracle fixtures already exist and are reused rather than re-derived:
`tests/vt/iq4nl_q5_0_golden_vectors.h` carries the pinned IQ4_NL/Q5_0 golden
vectors (real bytes of the shipped `per_layer_token_embd.weight` and the pinned
oracle's own `dequantize_row_iq4_nl` output) that pin the device-side op
bit-exactly. The host arm is held to the same vectors, so the two arms are
compared against one denominator, and the CPU per-row decode is the llama.cpp
`ggml_get_rows` behaviour this arm ports.

## Design

`EmbedGather(d, out, token_ids, table, vocab, H, what)` is the ONE call a dense
forward makes when it owns a device table. It owns both arms:

1. **Host arm** (`VT_HOST_EMBEDDING=1`): the table stays in `table.bytes` and is
never uploaded. A process-lifetime CPU queue runs `vt::Embedding` with a
`ViewOn` of the host bytes, which decodes one row per gathered id for every
table residency the loaders produce (bf16/f16/f32 and the GGUF block formats).
A bf16 table into a bf16 output takes a byte-copy fast path. The `[T,H]`
staging buffer is copied to `out`.
- The async runner may have spliced this step's sampled token into a
device-resident ids buffer, so the host arm consumes
`detail::TakeDeviceTokenIds()` and reads the ids back before the gather. The
splice runs through the same `detail::ApplyDeviceTokenIds` body the device
arm uses, so an override longer than the embed input is refused with the
caller's name instead of writing past the `[T]` buffer, and a SHORTER
override replaces exactly its prefix while the padded host tail stays.
- `VT_HOST_EMBED_TRACE=1` prints the arm and `T`.
- The flag is read ONCE per process (a function-static), so a serving process
cannot switch arms mid-run; a same-binary A/B is two processes.
2. **Device arm** (the shipped behaviour): `ResidentWeight` upload +
`ApplyDeviceTokenIds` + `vt::Embedding`, unchanged. Taken when the flag is off,
when the host bytes are gone (`bytes.empty()` or `host_released`), or when the
host arm declines for any reason; a one-shot `[host-embed] DISABLED` line names
the reason.

## Dependencies

- `vt::Embedding`'s CPU kernel and the CPU backend (in-tree, no new library).
- No GPU is needed for the host arm's correctness; a staging backend is needed to
exercise the `d_dev == nullptr` half.
- The async id override (`detail::TakeDeviceTokenIds`) must exist for the serving
loop; it is already in-tree from ENG-ASYNC-SCHED W4.

## Work breakdown

- **W1** (this PR): the seam, the flag, the CPU gather, the CPU-queue singleton,
the bf16 fast path, the async-override consumption, the wake-up of the dense
forward family, the docs, and the test.
- **W2** (owed): the end-to-end device-memory number and a decode-throughput A/B
on the target 24 GiB `sm_120a` card, both arms in one binary. Not a precondition
for the arm; the reason it is not in W1 is that a helper PR makes no speed claim.
- **W3** (later, if wanted): a config surface (`--offload-config`'s `vllm_cpp`
key) instead of an environment variable, if the operator prefers the config
document over env.

## Risks and decisions

| Risk / decision | Handling |
|---|---|
| A host gather adds a synchronize on the serving loop | The async runner's device-resident ids are consumed BEFORE the gather, so the host arm reads the spliced ids back rather than racing them; the decode-graph arms that hold ids on device keep the device gather and are listed in the header |
| A tied head keeps the table resident, so the memory saving is smaller | Recorded in the knob's documentation and in the header: the saving is the table's device residency for an untied embedding, and only the gather moves for a tied one |
| The flag is process-static, so a test binary cannot exercise both arms | The test binary enables it before `main` and is flag-ON by construction; the off arm is a plain early return and every other model suite runs it |
| `d_dev == nullptr` does not prove the arm ran on the CPU backend | `ResidentWeight` aliases host bytes when `is_cpu()`, so the test captures the one-shot `[host-embed]` banner and checks `HostEmbedInto`'s return value; the mutation that disables the arm makes 4 assertions fail |
| The host table's bytes are released by another path | The host arm declines on `bytes.empty()` / `host_released`, and `EmbedGather` then takes the device arm; the decline is a test case |
| The runner and the model disagree about this step's row count | The override's count is bounded against the embed input by `detail::ApplyDeviceTokenIds` on BOTH arms; a longer override throws with `what` naming the caller, a shorter one is the padded case and keeps the host upload's tail. Two test cases pin the boundary |

## Evidence

- The host arm's rows equal the pinned oracle through the same golden vectors the
device op is gated on; the arm is proven to have run (captured banner + return
value), not merely to have produced right numbers.
- The table is never uploaded on the host arm (`d_dev == nullptr`).
- The async override's shape is bounded: a longer-than-`T` override throws, and a
shorter one splices its prefix over the host upload while preserving the tail
(`tests/vllm/models/test_host_embedding.cpp`, the two override cases).
- `tests/vt/test_ops_embedding_quant` (6/6, 1637 assertions) is unchanged.

## Gates

- `ctest --test-dir build -R test_host_embedding`.
- `tests/vt/test_ops_embedding_quant` stays green (the device op is unchanged).
- `python3 scripts/check-env-doc.py` — the two new env vars are documented in
`docs/ENVIRONMENT.md` in this change (the checker's only remaining complaint is
the pre-existing `VT_VK_*` gap).
- The full `ctest --test-dir build`.
- `scripts/agent-preflight.sh --staged`.

## Owed

- The W2 measurement above (a 27B NVFP4/Q8mix GGUF on the 24 GiB `sm_120a` card,
the flag on and off, device memory and decode throughput), owed to the row.
- A dedicated config key (W3) only if the operator asks for one.
1 change: 1 addition & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -780,6 +780,7 @@ add_library(vllm STATIC
src/vllm/model_executor/models/qwen3_5_gguf_weights.cpp
src/vllm/model_executor/models/qwen3_gguf_weights.cpp
src/vllm/model_executor/models/qwen3_5.cpp
src/vllm/model_executor/models/host_embedding.cpp
src/vllm/model_executor/models/qwen3_5_common.cpp
src/vllm/model_executor/models/qwen3_5_dense.cpp
src/vllm/model_executor/models/qwen3_5_moe.cpp
Expand Down
2 changes: 2 additions & 0 deletions docs/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -233,6 +233,8 @@ portable/reference path. In normal operation leave them unset.
| `VT_ATTN_PREAMBLE_COOP` | off | `=1` selects the warp-per-item cooperative attention preamble arm (`AttnQkNormRopeGateCoopK`) on ROCm, mapping one warp per token item instead of the donor walk (`AttnQkNormRopeGateK`); read once per process like the sibling arms |
| `VT_GDN_COLPERM_KEEP_QUANT` | off | `=1` keeps the column-permuted `ssm_out`/`out_proj` tensor as Q5_K in tiled order (no `ReorderVCols`) and permutes the 4096-element GEMV input at runtime instead; the column reorder cuts across Q5_K block boundaries, so the weight cannot be permuted in place. Saves ~4x weight bandwidth (Q5_K ~5 MB vs bf16 20 MB per call) |
| `VT_GDN_ROWPERM_KEEP_QUANT` | off | `=1` keeps the row-permuted V-head GDN projections (in the tiled order the row permutation produces) as K-quant instead of expanding to bf16 at load; the runtime gather supplies the permutation. Opt-in; the default reorders then expands |
| `VT_HOST_EMBEDDING` | off | `=1` keeps the token table in host RAM and gathers the requested rows on a cached CPU queue, then copies the `[T,H]` result to the device. The CPU `vt::Embedding` kernel decodes ONE ROW per gathered id, so every table residency the loaders produce works (bf16/f16/f32 and the GGUF block formats); a bf16 table takes a byte-copy fast path. The async runner's device-resident id override is consumed before the gather. Falls back to the device gather when the flag is off or the host bytes are gone, with a one-shot `[host-embed] DISABLED` line naming the reason. Used by every dense forward that owns a device table (the Qwen3.5 family, MuseGlimmer, the shared Qwen3 dense driver, and the classic dense families: Gemma 1-4, GLM4, Granite, MiniCPM 1/3, OLMo2, OPT, Phi, Phi3, StableLM, Command-R, DeepSeek-V2, GLM-MoE-DSA, Dots3-Note, Nemotron-H, Voxtral); an untied embedding frees the table's device residency, a tied head keeps it for the GEMM. Decode-graph arms that hold their ids on device keep the device gather |
| `VT_HOST_EMBED_TRACE` | off | `=1` prints one `[host-embed] T=… override=…` line per host-side embedding gather, for telling which path a step took. Diagnostic only |
| `VT_GDN_SCAN_COOP` | off | `=1` selects the warp-per-row cooperative GDN scan arm (`GdnScanCoopK`) on ROCm, mapping one warp per output row instead of the donor walk (`GdnScanK`); read once per process like the sibling arms |
| `VT_GDN_PACKED_DECODE` | on (CUDA GDN) | Unpacked GDN decode path |
| `VT_GDN_DECODE_BV` | `32` (CUDA GDN decode experiment) | Exact `16` selects the byte-identical 16-value fused-recurrence tile; unset and every other spelling keep the 32-value schedule. Experimental opt-in; no release or cross-hardware default change |
Expand Down
43 changes: 43 additions & 0 deletions include/vllm/model_executor/models/host_embedding.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
// VT_HOST_EMBEDDING: gather the embedding rows on the CPU and copy the [T,H]
// result to the device, so the token table never occupies device memory. The
// CPU `vt::Embedding` kernel decodes ONE ROW per gathered id — the same per-row
// discipline as llama.cpp's ggml_get_rows — for every table residency the
// loaders produce: bf16/f16/f32 and the GGUF block-quant formats.
//
// Returns false (the caller uses the device path) when the flag is off or the
// table's host bytes are gone. A forward whose embedding is ALREADY a host
// gather does not need this; it exists for the forwards that own a device
// table. An untied embedding frees the table's whole device residency; a tied
// head keeps the table resident for its GEMM, so the flag then only moves the
// gather off the device.
//
// WHO CALLS IT. Every dense forward that owns a device table and takes its ids
// as a host vector calls `EmbedGather`. The sites that deliberately do NOT are
// the ones where a host gather cannot help or would add a synchronize the path
// exists to remove: decode-graph arms whose ids already live in a device tensor
// (qwen3_moe, gemma3, deepseek_v2, nemotron_h paged, qwen4_exp), the mm embed
// hooks whose ids arrive on device (qwen3_vl, dots3_note), the draft heads
// (qwen3_dflash/dspark), and the custom-residency loaders (kimi_linear).
#pragma once

#include <cstdint>
#include <vector>

#include "vllm/model_executor/models/dense_device_glue.h" // Dev, DBuf, OwnedTensor

namespace vllm {
namespace dense_attn {

bool HostEmbedInto(Dev d, DBuf& hidden, const std::vector<int32_t>& token_ids,
const OwnedTensor& table, int64_t vocab, int64_t H);

// The gather a forward should call when it owns a device table: the host arm
// above, else `ResidentWeight` + the async id override + `vt::Embedding`. The
// table and its shape are handed over UNRESOLVED so the upload happens only when
// the host arm declines. `what` names the caller in the override's shape check.
void EmbedGather(Dev d, DBuf& out, const std::vector<int32_t>& token_ids,
const OwnedTensor& table, int64_t vocab, int64_t H,
const char* what);

} // namespace dense_attn
} // namespace vllm
4 changes: 3 additions & 1 deletion scripts/check-agent-record.py
Original file line number Diff line number Diff line change
Expand Up @@ -389,7 +389,9 @@
# family, SERVE-RECIPE-ARGS / -REQUEST-LENGTH-GUARD, LOAD-GGUF-MMPROJ and the
# attention-window row. Bumped because a new row EXISTS, never to make a
# transition pass.
ENGINE_ROWS = 179
# 180 since 2026-09-29: +1 row (ENG-HOST-EMBEDDING, the host-resident token
# table). Bumped because a new row EXISTS, never to make a transition pass.
ENGINE_ROWS = 180

ENGINE_SUMMARY_SECTIONS = (
("Engine and scheduling", "Engine core and scheduling"),
Expand Down
6 changes: 3 additions & 3 deletions src/vllm/model_executor/models/commandr.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@

#include "vllm/model_executor/layers/linear.h" // UnquantizedMlpGateUpMethod seam
#include "vllm/model_executor/models/dense_attn_block.h" // shared device glue
#include "vllm/model_executor/models/host_embedding.h" // VT_HOST_EMBEDDING gather
#include "vllm/model_executor/models/device_pool.h" // DevicePool/Pool
#include "vllm/model_executor/models/qwen3_5_common.h" // HostLogits
#include "vt/backend.h"
Expand Down Expand Up @@ -194,9 +195,8 @@ DBuf ForwardBody(Dev d, const std::vector<int32_t>& token_ids,

DBuf hidden(d, DType::kBF16, {T, H});
{
Tensor dtab = ResidentWeight(d, weights.embed_tokens, {vocab, H});
DBuf dids(d, DType::kI32, {T}, token_ids.data());
vt::Embedding(d.q, hidden.t(), dtab, dids.t());
EmbedGather(d, hidden, token_ids, weights.embed_tokens, vocab, H,
"commandr embed");
}

StepInputs si = BuildStepInputs(d, positions, attn_meta, config);
Expand Down
Loading