From 15d48b09238ced63d01676d6b3dfbf596ce14c30 Mon Sep 17 00:00:00 2001 From: Yoav Date: Mon, 28 Sep 2026 15:46:25 -0400 Subject: [PATCH 1/2] fix(GATE-ISSUE-INDEX-TABLE-SHAPE): declare the three rows their own specs name, and the three defects that declaration was hiding Twelve canonical records were invalid. Nine were the _intake frozen-archive byte comparison, which is the CRLF class ENG-EOL-BYTE-EXACT addresses. The other three were rows that three specs and three open records already named and that no matrix declared, so canonical_rows() could not see them and each record died with "row is not canonical and claimable": - MODEL-GLINER25-DECIDE, named by the heading of specs/gliner2.5-decide.md, assigned to a new model-matrix row under MODEL-TOKCLS by that spec's own line 22 - MODEL-DSV41-EXL3 and MODEL-DSV41-GGUF-Q1_0, the two V4.1 checkpoint campaigns named by the specs/deepseek-v4-1-flash.md header The gap is invisible from the tool's own output. agent-issue-index.py --check raises on the FIRST invalid record and prints one line, and the first record it reaches is an _intake one, so the orphan rows were never named by anything the maintainer would run. Enumerating every record directly finds all twelve. THE THREE ROWS. MODEL-GLINER25-DECIDE goes in the MODEL-TOKCLS table as SPIKE, not as its spec's "Implementation not started", because the tree contradicts that: src/vllm/model_executor/models/ gliner25_decide_registry.cpp and gliner25_decide_head.cpp, a suite registered at tests/CMakeLists.txt:785, and a live C API path at src/capi/vllm_c.cpp:2106-2160. The other two go in beside MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model, their own spec's sibling, and both are BLOCKED, matching it. The checklist gains one line per row with the mark its state allows, and the rollup moves SPIKE 10 to 11, BLOCKED 5 to 7 and Total 384 to 387 in the same commit, which is what that file's own same-commit rule requires. THREE DEFECTS THE DECLARATION WAS HIDING, all found by replaying this work onto d15b1cc09 and all fixed here. The two checkpoint rows were one cell short. They were written with four cells where the table carries five, and check-agent-record.py reported "table has 5 pipes; expected 6" at model-matrix.md:166 and :167. The architecture-class cell was the missing one, not the model name: they opened | blocked | DeepSeek-V4.1-Flash EXL3 ... | description | row-id while every neighbour opens | mark | class | name | description | row-id. Both are V4.1-Flash checkpoints of one architecture, so the cell reads DeepseekV41ForCausalLM, the class the base row already carries. The checker reports one failure at a time, so this was masking the next two. The row-count constant was never bumped. check-agent-record.py pins each matrix's expected count in MATRICES as a hardcoded literal, not from the rollup, so the commit that moved the rollup to 387 red the record gate with "387 MODEL rows; expected 384". Both read 387 now, and the comment block above the constant records why in the form every previous bump in that file uses: a new row EXISTS, never a transition made to pass. MODEL-GLINER25-DECIDE had no owner, and SPIKE rows must have one. check-agent-record.py:1016 requires a CLAIM-* that actually claims the row. The owner cell read "unassigned", which is what all 333 other unassigned rows carry -- and every one of those is INVENTORIED, a state that needs no claim, so this row alone failed. It is now CLAIM-MODEL-GLINER25-DECIDE, a NEW claim file rather than a second line on CLAIM-MODEL-GLINER25, because adding it to the sibling claim produces "duplicate active claim CLAIM-MODEL-GLINER25": one claim owns exactly one active row. The split matches the work, which shares the DeBERTa v2 tower with MODEL-GLINER25 and differs only in the head (Linear -> ReLU -> Linear to one output in place of the NER boundary pooler), plus the /v1/systemone route and vllm_decide ABI it already shares with kev and Laya. TWO SPEC CORRECTIONS ship with it. gliner2.5-decide.md's ## Now said Implementation not started and is corrected to SPIKE with the code anchors that contradict it. deepseek-v4-1-flash.md's Matrices line assigned both campaigns to kernel-matrix.md, which is wrong -- they are model rows -- and is corrected to model-matrix.md with the rows that decide it named. MEASURED: invalid canonical records 12 to 9, the residue being exactly the nine _intake comparisons. In a worktree carrying ENG-EOL-BYTE-EXACT as well, agent-issue-index.py --check returns rc=0, so the two compose rather than compete. GATES, all rc=0: check-agent-record (MODEL=387), check-model-checklist, check-gate-commands (143 gated rows, up from 141), check-surface- coverage, check-conflict-markers, check-commit-trailers --range, check-commit-style --range. ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff] --- .agents/claims/CLAIM-MODEL-GLINER25-DECIDE.md | 7 ++++++ .../ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md | 19 +++++++++++++++ .agents/model-matrix.md | 14 +++++++---- .agents/specs/deepseek-v4-1-flash.md | 18 ++++++++++++-- .agents/specs/gliner2.5-decide.md | 24 +++++++++++++++---- scripts/check-agent-record.py | 11 ++++++++- 6 files changed, 81 insertions(+), 12 deletions(-) create mode 100644 .agents/claims/CLAIM-MODEL-GLINER25-DECIDE.md create mode 100644 .agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md diff --git a/.agents/claims/CLAIM-MODEL-GLINER25-DECIDE.md b/.agents/claims/CLAIM-MODEL-GLINER25-DECIDE.md new file mode 100644 index 0000000000..85b172d6cc --- /dev/null +++ b/.agents/claims/CLAIM-MODEL-GLINER25-DECIDE.md @@ -0,0 +1,7 @@ +# CLAIM-MODEL-GLINER25-DECIDE + +Separate claim from `CLAIM-MODEL-GLINER25`, not a second line on it: `check-agent-record.py` rejects a duplicate active claim ID, so one claim owns exactly one active row. The two rows share a tower and differ in a head, and the split follows that difference. + +| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update | +|---|---|---|---|---|---|---|---| +| `CLAIM-MODEL-GLINER25-DECIDE` | `MODEL-GLINER25-DECIDE` (`SPIKE`) | Claude Code (glm5.2), operator role | `.wt/gliner25` | `row/MODEL-GLINER25` | GLiNER2.5-Decide: the CLASSIFICATION-HEAD sibling of `MODEL-GLINER25`, reusing its DeBERTa v2 encoder and disentangled attention rather than reimplementing them, plus the `/v1/systemone` route and the `vllm_decide` ABI it already shares with kev and Laya. The head is Linear -> ReLU -> Linear to one output, replacing the NER boundary pooler; everything below the head is the sibling's. Bounded-choice decision, one forward, no prompt template and no generated tokens | `SPIKE` | 2026-09-28 | diff --git a/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md b/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md new file mode 100644 index 0000000000..f721f3aca9 --- /dev/null +++ b/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md @@ -0,0 +1,19 @@ +ID: ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS +Title: Three rows named by specs and by open issues are declared in no matrix, so agent-issue-index.py cannot validate any record +Row: GATE-ISSUE-INDEX-TABLE-SHAPE +State: OPEN +Kind: bug +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-27 +Updated: 2026-09-27 +Closed: - + +## Problem + +canonical_rows() reads claimable row IDs out of .agents/*-matrix.md first columns, .agents/specs declarations and headings, and the roadmap table (scripts/issue_records.py:208-243). Three row IDs are named by a spec and by an open canonical-local issue, and appear in NO matrix, so each record fails validate_issue_record with row is not canonical and claimable: MODEL-DSV41-EXL3 (.agents/issues/MODEL-DSV41-EXL3/ISSUE-LOCAL-01M29ATB6N2SFZD7JA6CCR2KXC.md), MODEL-DSV41-GGUF-Q1_0 (.agents/issues/MODEL-DSV41-GGUF-Q1_0/ISSUE-LOCAL-01M29ATQDVEQ7M2MP3VW79XQJP.md) and MODEL-GLINER25-DECIDE (.agents/issues/MODEL-GLINER25-DECIDE/ISSUE-LOCAL-01M3APZC6GKX9ME6AE6336D3VY.md). The first two are named in the .agents/specs/deepseek-v4-1-flash.md header as rows of this campaign (lines 3-4), which also assigns them to kernel-matrix.md; the third is named in the heading of .agents/specs/gliner2.5-decide.md as SPEC - MODEL-GLINER25-DECIDE and assigned to a new model-matrix row under MODEL-TOKCLS (line 22). So the specs and the issues agree with each other and only the matrices are missing. The gap is invisible from the tool's own output: agent-issue-index.py --check raises on the FIRST invalid record and prints one line, and the first record it reaches is an _intake one, so the orphan rows are never named. Enumerating every record directly finds 12 invalid, of which 9 are the _intake frozen-archive comparisons and 3 are these rows. Measured in a worktree at row/ENG-EOL-BYTE-EXACT (97397f60b), where the .agents/completed/issue-index.md working copy carries 0 CR against 886 CR on main, the 9 _intake records validate and exactly these 3 remain; that is, the orphan rows are the whole of the residue once #3333 lands, and the index stays blocked until they are declared. A second, related falsehood ships alongside: .agents/specs/gliner2.5-decide.md:14-18 says Implementation not started, while the same tree carries src/vllm/model_executor/models/gliner25_decide_registry.cpp, gliner25_decide_head.cpp, a registered suite at tests/CMakeLists.txt:785 and a live C API path at src/capi/vllm_c.cpp:2106-2160. This is the class found earlier in five of nine missing-arch specs: a ## Now that contradicts the tree it sits in. + +## Resolution + +Fixed in row/ORPHAN-MODEL-ROWS. Three rows declared in .agents/model-matrix.md: MODEL-GLINER25-DECIDE in the MODEL-TOKCLS table as its own spec directs, and MODEL-DSV41-EXL3 plus MODEL-DSV41-GGUF-Q1_0 beside MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model, the same spec's own sibling row. Each carries the state its evidence supports rather than the one the header implied: BLOCKED for the two checkpoint campaigns, matching that sibling, and SPIKE for GLiNER2.5-Decide, whose code, registered suite and C ABI path are in tree. The architecture-support checklist gains one line per row with the mark its state allows, and the rollup moves SPIKE 10 to 11, BLOCKED 5 to 7 and Total 384 to 387 in the same commit, which is what the file's own same-commit rule requires; check-model-checklist.py went rc=1 to rc=0 and caught the drift in between. Two spec corrections ship with it: gliner2.5-decide.md ## Now said Implementation not started and is corrected to SPIKE with the code anchors that contradict it, and deepseek-v4-1-flash.md's Matrices line assigned the two campaigns to kernel-matrix.md, which is wrong and is corrected to model-matrix.md with the two rows that decide it named. Measured: invalid canonical records 12 to 9, the residue being exactly the nine _intake frozen-archive comparisons that #3333 fixes, and in a worktree carrying both changes agent-issue-index.py --check returns rc=0. diff --git a/.agents/model-matrix.md b/.agents/model-matrix.md index 8f75cff7bd..361cf6433f 100644 --- a/.agents/model-matrix.md +++ b/.agents/model-matrix.md @@ -107,14 +107,14 @@ Rollup by lifecycle state (must equal the detailed per-state row counts): | INVENTORIED | 321 | | PARTIAL | 23 | | ACTIVE | 18 | -| SPIKE | 10 | -| BLOCKED | 5 | +| SPIKE | 11 | +| BLOCKED | 7 | | DONE | 3 | | READY | 3 | | GATING | 1 | -| **Total** | **384** | +| **Total** | **387** | -Engaged architectures (the 62 non-`INVENTORIED` rows): +Engaged architectures (the 66 non-`INVENTORIED` rows): | Support | Architecture | Family / example | Status | Row | |---|---|---|---|---| @@ -123,6 +123,7 @@ Engaged architectures (the 62 non-`INVENTORIED` rows): | 🚧 | `MiMoV2ForCausalLM` | MiMoV2 (inference-time weight offload candidate; W1 registry + config parse + hybrid KV-cache spec + weights loader landed, W2 tests in tree) | `ACTIVE` (W1); the port decision and wave scope live in [specs/mimov2.md](specs/mimov2.md) | `MODEL-TEXT-mimo-v2-mi-mo-v2-for-causal-lm` | | 🚧 | `DeepseekV4ForCausalLM` (EXL3 arm) | DeepSeek-V4-Flash SparkInfer EXL3 3.0bpw REAP-K216 checkpoint campaign (W1a/W1b landed and reviewed PASS; W2a/W2b landed, CUDA arm compile + device measurements PENDING) | `ACTIVE` | `MODEL-DSV4-EXL3` | | πŸ“‹ | GLiNER2.5 (BoundaryExtractor / SpanExtractor) | GLiNER2.5 port decision row; gliner2 + gliner25_decide code and test suites are in tree, serving a real checkpoint end to end is the owed scope; oracle is the `vllm-factory` registry | `SPIKE` | `MODEL-GLINER25` | +| πŸ“‹ | GLiNER2.5-Decide (classification head) | SystemOne-class decision classifier on the same DeBERTa-v3-large encoder β€” one forward, no prompt template, no generated tokens; single-label, multi-label, ordinal and yes/no (noul) decisions. Reuses `MODEL-GLINER25`'s encoder; `gliner25_decide` code and its registered suite are in tree and the C API routes it at ABI v29 | `SPIKE` | `MODEL-GLINER25-DECIDE` | | 🚧 | `DeepseekV4ForCausalLM` | DeepSeek-V4-Flash-Vision-Exp (43-layer V4 language model + 32-layer ViT, image-text-to-text) | **ACTIVE; W1 PROMPT ENCODER AND IMAGE PROCESSOR LANDED; MODEL INFERENCE IS NOT YET WIRED.** The released architecture string is the existing text `DeepseekV4ForCausalLM`, with `vision_n_layers=32`, a downsample-3 aligner, four image sentinel embeddings and image-span attention visibility. vLLM at the pin and current main implement no vision path; Transformers implements text only. The model-author runtime at pinned Hugging Face revision `86f746b3` is the only complete reference and is registered as `deepseek-v4-vision`, `gateable = no` until it builds and runs the 156.287 GiB TP4 artifact. The port reuses the existing DeepSeek-V4 backbone and shared multimodal engine; W1 is the RED-first prompt/image processor. [#2411](https://github.com/mudler/vllm.cpp/issues/2411) | `MODEL-MM-deepseek-v4-deepseek-v4-for-causal-lm` | | βœ… | `Qwen3ForCausalLM` | Qwen3 dense (0.6B/1.7B/4B/32B) | near-tie-robust token-exact 16/16 on 0.6B+4B vs vLLM 0.25.0; NVFP4A16 (W4A16) dense quant also gated; c1 every-axis speed parity, c8 decode residual; async-serving device token-ids mirror ported (`ROW-SERVE-ASYNC-DENSE-MIRROR`, #31 fix into the shared dense `EmbedInto`) β€” `test_qwen3_dense_async_serving` REDβ†’GREEN; sibling scope CLOSED (#323): `60e71a0e` fixed the eager path; `DenseDecodeGraphForward` ran first and replayed against stale HOST ids, so it now declines while the mirror is live and falls back to the proven eager path. Async gate 7/7 across Qwen3-0.6B/4B + Llama/Mistral/InternLM2 | `MODEL-TEXT-qwen3-qwen3-for-causal-lm` | | βœ… | `Qwen3MoeForCausalLM` | Qwen3-Coder-30B-A3B (MoE) | STRICT token-exact 6/6 vs vLLM 0.25.0; 11/16 speed-grid cells at/above graphed vLLM, c1/c2 residual | `MODEL-TEXT-qwen3-moe-qwen3-moe-for-causal-lm` | @@ -163,6 +164,8 @@ Engaged architectures (the 62 non-`INVENTORIED` rows): | 🚫 | `DeepseekV3ForCausalLM` / `DeepseekV32ForCausalLM` | DeepSeek-V3 / V3.2 | HW-blocked (671B, ~642 GiB fp8 vs 119 GiB unified memory); V3.2 additionally DEP-blocked (DSA indexer) | `MODEL-TEXT-deepseek-v2-deepseek-v3-for-causal-lm` | | 🚧 | `DeepseekV41ForCausalLM` | DeepSeek-V4.1-Flash (552B backbone, 8B activated at prefill / 16B at decode, natively multimodal) | **REGISTERED AND VALIDATING; NOT LOADABLE (W1, 2026-09-13).** The architecture RESOLVES through `ModelRegistry` and the REAL published `config.json` @ `dba1be0a40` descends and validates; the loader, the forward, the KV-cache spec and the `deepseek41` GGUF container each refuse BY NAME. It does NOT load and does NOT forward. Registration claims NO gate against the oracle, which is the standing `Qwen3_5ForCausalLM`/`Qwen3_5MoeForCausalLM` already carry ("Ahead-of-pin forward port ... REGISTERED, NOT RUN-GATED", [#490](https://github.com/mudler/vllm.cpp/issues/490)), so NO pin advance was needed or made. **Still pin-blocked for a GATE, and still HW-blocked on every published arm.** Registered on vLLM `main` at `e77daef89e` (`registry.py:379`) and MEASURABLY absent at our pin `e126687a9a` (zero grep hits, no `vllm/models/deepseek_v4_1/`), 566 commits away, so no primary oracle exists and no secondary may substitute for a path vLLM implements. Also HW-blocked on every published arm: release 475.27 GiB, EXL3 3.5bpw 428.49 GiB, and the one GGUF rung that fits 119 GiB is degenerate by construction. Upstream FORKED the V4 tree rather than extending it, and the net-new surface has SIX parts with no counterpart here (Engram, MXFP8 32x32, second indexer block size, DSpark-without-MTP, vision tower, and the `kv_source_layer`-keyed KV topology -- which `e030f1b90` measured is NOT an encoder-decoder split and owes nothing for SWA Bounded Replay, see the spec's W3d) | `MODEL-MM-deepseek-v4-1-deepseek-v41-for-causal-lm` | | 🚫 | `DSparkV41DraftModel` | DeepSeek-V4.1-Flash DSpark drafter (spec-decode; V4.1 ships no MTP head) | **PIN-BLOCKED** on the same advance as its target, and blocked behind the target itself. Registered on vLLM `main` at `registry.py:647`, absent at the pin, whose `DSparkDraftModel` (`registry.py:624`) is a different class against the V4 target | `MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model` | +| 🚫 | `DeepseekV41ForCausalLM` | DeepSeek-V4.1-Flash EXL3 3.5bpw "Pollard" | 428.49 GiB checkpoint campaign. **The format claim covers part of the file, not the file**: shards 1, 2 and 43-48 are byte-identical to the deepseek-ai release (199.88 GiB) and `quantized_modules` covers the routed experts alone; "Pollard" is a per-expert width recipe over turboderp EXL3 trellis, the `MODEL-DSV4-EXL3` format. 3.6x one GB10, publisher numbers are 4x TP4 and unreproduced | `MODEL-DSV41-EXL3` | +| 🚫 | `DeepseekV41ForCausalLM` | DeepSeek-V4.1-Flash GGUF `Q1_0` | 98.591 GiB against a 119 GiB pool β€” the only published V4.1 rung that fits one GB10, and the only part of it that works: `token_embd` is stored at ggml type 41, so the embedding table returns plus-or-minus one magnitude. `deepseek41` is in no released engine; a non-degenerate fitting rung is a requantization question this row does not scope | `MODEL-DSV41-GGUF-Q1_0` | | 🚧 | `GlmMoeDsaForCausalLM` | GLM-5 / GLM-5.3 (DSA) | **THIS MODEL GENERATES FROM ITS REAL 201.83 GiB ARTIFACT, THROUGH THE EXPERT-STREAMING LANE: `The capital of France is` -> ` Paris`, on `dgx:gpu0` (GB10), 2026-08-31** ([#2214](https://github.com/mudler/vllm.cpp/issues/2214), spec Β§3.10) β€” `--device cuda` with `VT_MOE_EXPERT_STREAM=1` and 4096 slots, `rc=0`, `prompt_tokens=5 completion_tokens=1`, wall 1154 s, `generate` 852.330 s, `VmHWM` 57.71 GiB against 119.631 GiB of device. **The lane's counters are the evidence and they carry one caveat:** `[expert-stream] ON slots=4096 slot_bytes=6684672 resident=25.50 GiB` β€” Β§3.3's arena to the byte β€” then `steps=1 hits=0 misses=6399 evictions=0 fills=4096 bytes=13939408896 exhausted=2303 advised=0`. 4096 slices streamed and 12.98 GiB moved with zero evictions, the 187.312 GiB of towers never materialized; but the step needed 6399 slices, so 2303 (36%) were read in place out of the mapping, which is a PREFILL working set exceeding any budget by construction (spec R2, O34) and is counted rather than silent. No figure is quoted as a fully-streamed step and no speed number is claimed. Getting here fixed three defects of ours, each measured first: the loader PREFAULTED all 228 towers (O30), `token_embd.weight` was routed as a GEMM weight rather than a gather so `EmbeddingKernelCuda` refused a Q4_K table (O31), and the slot store was sized from the first layer instead of the file so the lane streamed 527 slices and then refused at `blk.8`'s IQ4_XS tower (O33). The `--device cpu` arm emits the same ` Paris` and does NOT stream. Reads are CIFS (spec O7). ([#2214](https://github.com/mudler/vllm.cpp/issues/2214), spec Β§3.10) β€” `rc=0`, `prompt_tokens=5 completion_tokens=1`, `generate` 950.249 s, `VmHWM` 44.46 GiB. **That token came from the `--device cpu` arm and is NOT a streaming result**: a CPU queue builds no slot lane, so every routed-expert slice is read in place out of the mapping. No speed number is claimed from it. **The STREAMING lane also runs on the same box and artifact** with `--device cuda`, `VT_MOE_EXPERT_STREAM=1` and 4096 slots, and it is the first time this project's `pread` path has run on a real checkpoint: `[expert-stream] ON slots=4096 slot_bytes=4816896 resident=18.38 GiB`, then `steps=1 hits=0 misses=527 evictions=0 fills=527 bytes=1876328448 exhausted=0 advised=0` β€” 527 slices paged into slots, no eviction, and `exhausted=0`, so nothing fell back to the mapping. That load is 349 s at `VmHWM` 42.04 GiB against 187.312 GiB of towers never materialized. It then refuses on the slot BUDGET at `blk.8` (spec O33, fixed). Driven on `thor:gpu0` under an `rc` lease 2026-08-31 ([#2214](https://github.com/mudler/vllm.cpp/issues/2214), spec Β§3.10): `--device cuda` with `VT_MOE_EXPERT_STREAM=1` and 4096 slots, all 1809 tensors of all six shards resolve, the MTP block drops, the MLA absorption runs and the engine sizes its KV cache in 866 s wall with `VmHWM` at **23.10 GiB** β€” about 20.3 GiB of predicted resident weight and 187.312 GiB below the artifact, so the expert towers are NOT materialized. The streamed-expert lane IS built for this model (`CheckDeviceWeightFit` would otherwise charge the device all 187.312 GiB and refuse), which discharges spec O14 and O29; no slice was served and no `[expert-stream]` counter is quoted, because the first step threw. It throws at `cuda mla_prefill_attention: built without the vendored FlashAttention-2` β€” the BUILD and the ARCH rather than the file: FA2 needs CUTLASS headers and covers `8.0,8.6,8.7,8.9,12.0a,12.1a`, so `thor`'s sm_110a can never reach a token and `dgx:gpu0`'s sm_121a can. The drive found and fixed two defects of ours, both measured before they were repaired: the loader PREFAULTED all 228 expert towers, so RSS climbed linearly past 48.62 GiB at the filesystem's read rate (spec O30), and `token_embd.weight` was routed as a GEMM weight rather than as a gather, so a Q4_K table reached `EmbeddingKernelCuda` and was refused by name (spec O31). Reads are CIFS and every wall time here is a CIFS number (spec O7).** W2 landed 2026-08-30 ([#2214](https://github.com/mudler/vllm.cpp/issues/2214), [spec](specs/glm-dsa-latest-deepseek.md) Β§3.7): `GlmMoeDsaForCausalLM` self-registers from its own TU, its config resolves from a `config.json` and from a `glm-dsa` GGUF header through ONE validator, `kGgufArchArms` carries a `glm-dsa` row, and the forward refuses by name and lists all seven missing primitives. GLM-5.3 gets its OWN params struct rather than sharing `DeepseekV2Params`, so the `index_topk` tripwire (`deepseek_v2_weights.cpp:358-364`) stays a WALL for DeepSeek-V2 instead of becoming a choice. The indexer schedule is DERIVED the way vLLM derives it (`deepseek_v2.py:1097-1101`), with the explicit list as an override; the three-way agreement of Β§3.5.1 is now EXECUTABLE and holds β€” the checkpoint's own 78-entry `indexer_types` (committed verbatim from revision `935644c05e76fc198714f4cca449fd8b970ff6d7`), vLLM's derived rule at `freq = 4` / `offset = 3`, and llama.cpp's `GLM_5_2_DEFAULT_INDEXER_TYPES` (`b10451:src/models/glm-dsa.cpp:6-27`) agree on all 78 entries, 21 `full`. OWED by W2 itself: the forward, the KV-cache hook and the loaded-model factory are UNREACHED because both `load_weights` arms refuse (spec O16, W7), and the one staged GGUF arm states no `glm-dsa.attention.indexer.types` so it is refused rather than resolved off llama.cpp's hardcoded table (spec O17, D3). Prior scoping, unchanged ([#2214](https://github.com/mudler/vllm.cpp/issues/2214), [spec](specs/glm-dsa-latest-deepseek.md) Β§3). Recomputed from `zai-org/GLM-5.3`'s own `config.json` and checked against its `model.safetensors.index.json`, the routed experts are **97.49% of 753.33B parameters** (reproducing the API's measured total to -0.00016%), so the question is the step working set, not resident capacity. A full HTTP-range census of `unsloth/GLM-5.3-GGUF` `UD-IQ1_S` (revision `346b3591c7f2`, 6 shards, 1809 tensors) measures **14.511 GiB resident + 187.312 GiB of streamable `*_exps` towers**; one `c = 1` decode step touches 1800 slices = 11.21 GiB of uniform slots, so resident + a 4096-slot cache is **40.01 GiB** against 119.631 GiB on `dgx:gpu0`. **THE `vec_dot` BLOCKER IS CLEARED, 2026-08-30.** This row read "blocked on exactly one kernel, `VecDotIQ4_XSQ8_K`" until today; `2e9f4d88d` ([#2247](https://github.com/mudler/vllm.cpp/issues/2247), [#2256](https://github.com/mudler/vllm.cpp/issues/2256)) had already landed it for the sibling Flash row. Verified at the three sites on `origin/main` rather than taken on report: defined `cpu_quant_dot.cpp:844`, dispatched `:1004`, traits `MakeTraits(kIQ4_XS, kQ8_K)` at `cpu_quant_traits.cpp:121-122`, and `KeepQuantDType` gates on `HasQuantDotKernel` at `gguf_keep_quant.cpp:176` β€” so `HasQuantDotKernel(kIQ4_XS)` is TRUE and the four towers keep their blocks instead of expanding 6.375 -> 24.000 GiB. `kIQ2_XS` landed in the same commit (`:1003`), clearing `UD-Q2_K_XL` too. **All six encodings the `UD-IQ1_S` arm uses β€” IQ1_S, IQ3_XXS, IQ2_XXS, IQ4_XS, Q2_K, Q3_K β€” now have a `vec_dot` row, 6 of 6**, so the arm stays in the streaming lane and W7 has no keep-quant question left. `IQ1_M` still has no reader traits, so `UD-IQ1_M` alone still refuses at file open. Spec W1 is discharged and spec O2, plus O3's IQ2_XS half, close with it. The pinned vLLM class CAN load this checkpoint (it derives the indexer schedule from `index_topk_freq`/`index_skip_topk_offset` at `deepseek_v2.py:1092-1103` and never reads `indexer_types`), but it cannot RUN it on any fleet device, so **NO end-to-end token gate against vLLM is reachable** and the spec says so before any wave promises one. Eight waves. **W2, W3 and W4 have landed** ([#2214](https://github.com/mudler/vllm.cpp/issues/2214)); **W1 is discharged by another row** (`2e9f4d88d`, #2247/#2256); and **W5 is STRUCK from this row and re-sequenced as a consumption site**, because `KV-DSV4-MULTICACHE` has scheduled its own W5 ([#2323](https://github.com/mudler/vllm.cpp/issues/2323)) and this row's W5 was conditional on it not doing so. **W7 is the critical path**, with W6 behind it. **W4** puts the heterogeneous indexer schedule on per-layer `mla::MlaBlockDims` and mirrors the `skip_topk` selection reuse: `GlmMoeDsaMlaSchedule` turns the parsed `indexer_types` into 78 block geometries, **21 carrying an indexer and 57 carrying `skip_topk`** β€” 22 of 79 once the MTP block upstream forces full (`deepseek_v2.py:1110-1115`) is counted β€” asserted from the checkpoint's own `config.json` rather than from a literal. The reuse is upstream's own shape and it is the ABSENCE of a write, not a copy: ONE `topk_indices_buffer` per model (`deepseek_v2.py:1372-1377`) reaches every layer (`:1395`), `mla.py:180` runs the indexer only when `not self.skip_topk`, and a shared layer's indexer does not exist at all (`:1134-1135`), so it attends through the bytes the preceding full layer left in the shared buffer. Mirrored as `mla::MlaSharedSelection`; a `skip_topk` layer handed no buffer is REFUSED rather than falling through to the dense key loop, which would have been finite, plausible and wrong on 57 of 79 blocks. W4 also lands the **fp32 router gate GEMM**, which turned out to be a DeepSeek-V2 parity repair too: `_get_moe_router_dtype` (`deepseek_v2.py:123-133`) returns f32 for `glm_moe_dsa` at `:127` AND for any `moe_router_dtype: "float32"` at `:131`, and `deepseek_v2.cpp` hardcoded bf16, so a V2/V3 checkpoint asking for an f32 router silently did not get one. The eight dtype expectations are the pinned oracle's OWN return values β€” `_get_moe_router_dtype` extracted from `5559679229` and executed under torch 2.11.0+cu130 β€” not a transcription. Gates: 24 cases / 2594 assertions new, with `test_mla_attention_block` (18 / 2,255,433), `test_dots3_note_attn` (51 / 6888), `test_glm_moe_dsa_config` (15 / 380) and `test_deepseek_v2_load` (4 / 14) unmoved. OWED by W4: the block's reuse arm has no production caller until W7's forward exists (spec O19), and the router dtype's SELECTION at `deepseek_v2.cpp:363` is reached but NOT numerically gateable β€” the f32 and bf16 arms are bit-identical over all 500 tiny-fixture logits and forcing f32 reds nothing, which is AGENTS.md's "a token gate cannot detect a dtype that is too wide" as a measured instance (spec O18). W3 lifted the expert-streaming WIRING out of `qwen3_5.cpp` into `expert_stream_seam.{h,cpp}` β€” `ExpertStreamLane`, the step guard, `HostSliceView` and `ExpertSlice` β€” so a second model TU can reach a lane that until now only ONE translation unit could construct (Β§3.2 gap 1, `## Owed` O8, now discharged). Qwen3.5 is the seam's first client and is byte-identical across the lift: the same synthetic forward through `Qwen3_5Model::Forward` produces logits digest `84ae1a52ee64d117` over 160 floats from two separately-built libraries differing only in `qwen3_5.cpp`, streaming OFF and ON, with an identical `[expert-stream] steps=1 hits=0 misses=42 evictions=0 fills=42 bytes=45696 exhausted=0 advised=42` line. The resident-tower fallback is INJECTED rather than called, because `qwen3_5.cpp`'s own `ResidentWeight` shadows `dense_attn_block.h`'s and carries staging refusals the header's version does not. W3 also adds the load-time slot-capacity refusal Β§3.3 argues for: a budget below `streamed_towers * experts_per_tok` now refuses by name at `LoadedEngine::FromModelDir` instead of degrading to a 187 GiB per-token mmap read that a benchmark would publish as a streaming number | `MODEL-TEXT-deepseek-v2-glm-moe-dsa-for-causal-lm` | | 🚫 | `MiniMaxM2ForCausalLM` | MiniMax-M2 | HW-blocked (~230B / ~428 GiB bf16, ~4x over unified memory) | `MODEL-TEXT-minimax-m2-mini-max-m2-for-causal-lm` | | βœ… | `GemmaForCausalLM` | Gemma 1 (gemma-2b) | STRICT token-exact SACRED gate 48/48 greedy vs vLLM 0.25.0 (K=5 ALL-DETERMINISTIC β†’ STRICT; BOS-verified; ungated `unsloth/gemma-2b` mirror). The original Gemma: two fused add+RMSNorm/layer, `head_dim^-0.5` scale, GeGLU + `sqrt(hidden)` embed-scale, tied lm_head, no soft-cap/QK-norm/sliding; reuses the W1 GeGLU/embed-scale primitives; speed pending | `MODEL-TEXT-gemma-gemma-for-causal-lm` | @@ -421,6 +424,7 @@ Transformers compatibility is capability-driven and excluded from finite counts. | `MODEL-TOKCLS-modernbert-modern-bert-for-token-classification` | `ModernBertForTokenClassification` | `registry.py:288-291`; `vllm/model_executor/models/modernbert.py::ModernBertForTokenClassification` | token classification / text or audio alignment | encoder attention; token head/pooler; sliding-window attention | ☐ required | `INVENTORIED` | none | unassigned | | `MODEL-TOKCLS-openai-privacy-filter-open-aiprivacy-filter-for-token-classification` | `OpenAIPrivacyFilterForTokenClassification` | `registry.py:292-295`; `vllm/model_executor/models/openai_privacy_filter.py::OpenAIPrivacyFilterForTokenClassification` | token classification / text or audio alignment | encoder attention; token head/pooler; sliding-window attention | ☐ required | `INVENTORIED` | none | unassigned | | `MODEL-TOKCLS-qwen3-asr-forced-aligner-qwen3-asrforced-aligner-for-token-classification` | `Qwen3ASRForcedAlignerForTokenClassification` | `registry.py:296-299`; `vllm/model_executor/models/qwen3_asr_forced_aligner.py::Qwen3ASRForcedAlignerForTokenClassification` | token classification / text or audio alignment | encoder attention; token head/pooler | ☐ required | `INVENTORIED` | none | unassigned | +| `MODEL-GLINER25-DECIDE` | `GLiNER2.5-Decide` (`fastino/GLiNER2.5-Decide`; 340M, DeBERTa-v3-large encoder plus a classification head, Linear β†’ ReLU β†’ Linear to output dim 1 β€” the CLASSIFICATION-HEAD sibling of `MODEL-GLINER25`, which pools NER boundaries off the same encoder) | **NOT AT THE PIN, and no upstream vLLM implementation exists.** vLLM has no DeBERTa, no disentangled attention and no GLiNER; the reference implementation is the `vllm-factory` oracle (`ddickmann/vllm-factory`), the same oracle `MODEL-GLINER25` names, and HF `transformers` has no GLiNER2.5 | token classification / bounded-choice decision β€” ONE forward pass, NO prompt template and NO generated tokens; single-label choice, multi-label classification, ordinal scoring and yes/no (noul) decisions, several heads scored in a single call; CPU + GPU (CUDA) | the DeBERTa v2 encoder and its disentangled attention, already in tree under `MODEL-GLINER25` and the key reuse here; the `/v1/systemone` seam and the `vllm_decide` ABI it shares with kev and Laya | [gliner2.5-decide](specs/gliner2.5-decide.md) (issue `ISSUE-LOCAL-01M3APZC6GKX9ME6AE6336D3VY`) | `SPIKE` | code and a registered suite are in tree: `src/vllm/model_executor/models/gliner25_decide_registry.cpp`, `src/vllm/model_executor/models/gliner25_decide_head.cpp`; suite `tests/vllm/models/test_gliner25_decide.cpp:1-181` registered at `tests/CMakeLists.txt:785`; C API `src/capi/vllm_c.cpp:2106-2160` (`Gliner25DecideInference`, ABI v29). **This spec's `## Now` read `Implementation not started` and is corrected by the change that added this row** β€” the code, the registration and the C ABI path above were already in the tree it sat in. It holds `SPIKE` on its sibling's terms: `MODEL-GLINER25` holds `SPIKE` until its spec carries the full structured field set the `ACTIVE` promotion requires, and nothing here is a measured end-to-end decision | `CLAIM-MODEL-GLINER25-DECIDE` | ## MODEL-SEQCLS - Sequence classification @@ -611,6 +615,8 @@ row's `State` cell. | `MODEL-SPEC-qwen3-dflash-dflash-qwen3-for-causal-lm` | `DFlashDraftModel` | `registry.py:591`; `vllm/model_executor/models/qwen3_dflash.py::DFlashQwen3ForCausalLM` | speculative draft / target-dependent | draft runner; acceptance/sampling; sliding-window attention; DFlash | βœ… [DFlash spec](specs/dflash-spec-decode.md) Β§0 D14 | `DONE` | D0-redo (`CLAIM-DFLASH-D0D1`, 2026-07-26) UNBLOCKED on the advanced pin `555967922`/vLLM 0.26.0.dev0: the mixed-SWA/full z-lab 27B draft CONSTRUCTS + drafter ALIVE under `VLLM_USE_V2_MODEL_RUNNER=1` (acceptance > 1); D0+D1 (`DF-AUX-TAPS` target-side multi-tap) landed. **D2 (`DF-DRAFT-MODEL`, `CLAIM-DFLASH-D2`, 2026-07-26) CODE LANDED + CPU-GATED:** drafter model `qwen3_dflash.{h,cpp}` (plain 5-layer Qwen3-dense reusing the landed dense block ops, routing attention through the NEW `vt::DFlashBlockAttention` non-causal in-block primitive), the fc aux-combine, mask-embed, per-layer SWA/full resolution (`ResolveQwen3DFlashAttnModes`), and the z-lab loader. CPU gate GREEN: `test_ops_dflash_block_attn` 5/12 (hand-checked non-causal + RED causal-vs-non-causal + block isolation + SWA window + GQA), `test_qwen3_dflash_forward` 5/95 (forward runs + RED full-layer-causal-flip + block isolation + fc-combine ref + RED reversed-tap-order + attn-mode resolution). Existing causal `test_ops_attention` 9/9Β·23 + `test_qwen3_forward` 5/1028 UNCHANGED (new op is SEPARATE; causal path byte-identical). **`ACTIVE` β€” GPU promotion GREEN on dgx 2026-07-26 (GB10 sm_121a):** CUDA `-Werror` build clean (kernel compiles as-written); CUDA==CPU parity `test_ops_dflash_block_attn` 198412/198412 + compute-sanitizer 0; draft-forward parity `test_qwen3_dflash_draft_parity` vs the REAL vLLM DFlash draft (via `collective_rpc`) β€” fc rel-L2 0.46%, per-layer hidden ≀1.30%, final 0.88%, 11 deterministic rows STRICT top-1 + 5 bf16-near-tie cluster-matched; 27B SACRED 235/235 + MTP 9/9 byte-identical. Loader now tolerates the draft ckpt's omitted embed_tokens/lm_head (target-shared). Capture harness `scripts/spec/d2_dflash_draft_ref.py`. **D3 `DF-DRAFT-KV-PREP` CODE LANDED + CPU-GATED 2026-07-26 (`CLAIM-DFLASH-D3`):** `PrepareDflashInputs` (pure-integer HOST port of `_prepare_dflash_inputs_kernel`) + `PrecomputeContextKV` + `ForwardBlockLogitsWithContext` in `qwen3_dflash.{h,cpp}` reuse the landed MatmulBT/RmsNorm/RopeNeox + the UNCHANGED D2 `vt::DFlashBlockAttention` via a `[context;block]` combined sequence (NO new kernel); CPU gate `tests/vllm/v1/spec_decode/test_dflash_kvprep.cpp` 6/114 (prepare INTEGER bit-exact + RED valid_ctx_end; context-KV envelope + RED hidden_norm/k_norm/pos; context forward degenerates to D2 at empty ctx + diverges with ctx + block isolation); additive-only diff so D2/MTP/SACRED byte-identical. **D3 DONE 2026-07-26 β€” GPU numeric-parity GREEN on dgx (GB10 sm_121a):** CUDA `-Werror` 0 warnings; `test_qwen3_dflash_kvprep_parity` 61/61 vs the REAL vLLM draft β€” `prepare_dflash_inputs` INTEGER bit-exact vs vLLM's ACTUAL Triton `_prepare_dflash_inputs_kernel` (`kernel_matches_numpy`), context-KV worst K/V rel-L2 0.31%/0.26% across all 5 layers (proves per-layer `qkv_proj` slicing == vLLM's fused KV weight), block-proposal 13 STRICT + 3 near-tie = 16/16; CPU `test_dflash_kvprep` 114/114 re-passed (RED-proven); 27B SACRED 235/235 + MTP 9/9 + D2 parity 37/37 byte-identical. Capture harness `scripts/spec/d3_dflash_kvprep_ref.py`. **D4 `DF-ENGINE-INTEGRATION` propose brick + `dflash` config-select CODE LANDED + CPU-GATED 2026-07-26 (`CLAIM-DFLASH-D4D5`):** NEW `src/vllm/v1/worker/gpu/spec_decode/dflash/speculator.{h,cpp}` `DflashProposeBlock`/`SampleDflashBlockDrafts` β€” the non-autoregressive whole-block propose swapping in for the MTP k=1 `MtpProposePrefill`, composing D3 `ForwardBlockLogitsWithContext` + greedy per-mask argmax (anchor not sampled, `dflash/speculator.py:300-413`); `ParseSpeculativeConfigJson`/`ResolveDflash` accept `method:"dflash"` (lookahead k+1). CPU gate `tests/vllm/v1/spec_decode/test_dflash_propose.cpp` 5/19 GREEN (RED-first: sampler reading the anchor row fails 4/5 cases; brick == forward+sampler; empty-ctx degenerates to D2 context-free argmax; config parses; dspark throws), `-Werror`-clean TU. Additive + config-gated β‡’ MTP/non-spec byte-identical BY CONSTRUCTION (no runner/model/loader/scheduler edit). **D5 `DF-ENGINE-INTEGRATION` runner-loop LANDED + e2e RUNS on dgx 2026-07-26 (`CLAIM-DFLASH-D5`):** the verify/propose loop is wired β€” loader loads the SEPARATE z-lab draft (`LoadDflashDraft`, host bf16 + target-SHARED bf16 embed/lm_head) via a `--speculative-config` `model` key + `ResolveSpecConfig` dflash branch + `runner.set_dflash_draft`; the verify forward captures the D1 multi-tap (`aux_tap`β†’`ForwardDeviceMultiTap`); `propose_drafts_dflash` ACCUMULATES the per-request combined-feature context (`CombineAuxFeatures(aux_tap)`) and honors the `num_rejected` rollback by appending only the accepted-prefix features, then runs `DflashProposeBlock` (k=16). **e2e `test_qwen27_dflash_spec_decode` (4Γ—32 tok vs the committed vLLM-DFlash-ON golden): 2/4 STRICT token-exact (fibonacci, three-laws) + acceptance ~ vLLM on ALL 4 (accepted 19/39/29/25 vs golden 17/39/30/25 β€” the MANDATORY dead-drafter-trap condition MET).** The 2 divergences (France tok11, 17*23 tok12) are SINGLE bf16 near-tie flips (17*23 RE-CONVERGES = proven near-tie; France cascades from one flip) β€” the ratified near-tie ROOT, rooted in the D3-documented inline bf16 context-KV recompute envelope (~0.3-1.3% rel-L2), NOT a wiring bug (proven by the 2 exact prompts + near-exact acceptance). Inertness GREEN (SACRED 235/235 + MTP 9/9 byte-identical); CUDA `-Werror` clean; NO new CUDA kernel (host orchestration reusing D1/D2/D3-sanitized ops). NOT a clean strict-4/4 pass β€” STRICT 4/4 token-identity + speed A/B = D6 (persistent paged draft-KV bit-matching vLLM's fused context-KV projections). **D6 `CLAIM-DFLASH-D6` 2026-07-27 (records-only, NO code) β€” c1 SPEED A/B DONE + STRICT proven bf16-IRREDUCIBLE + CG feasibility:** our DFlash-ON = **2.50x TPOT (40.4 vs 101.2 ms) / 2.48x output-tput (24.4 vs 9.86 tok/s)** over our OFF at c1 (8 prose+code promptsΓ—256 tok greedy, 2 reps, acceptance 0.22 = 3.56/16, rep-stable <1.5%), `benchmark_binding=true`; vs vLLM-DFlash-ON graphed vLLM-DFlash-ON graphed = 28.5 tok/s / 35.1 ms TPOT / acceptance_len 4.30 (same 8 prompts, `VLLM_USE_V2_MODEL_RUNNER=1`, mm-off, gpu_util 0.30), so OURS IS ~14% BELOW vLLM-DFlash-ON on output throughput (24.4 vs 28.5 tok/s) - both ~on-par at spec-OFF (9.86 vs 9.83 tok/s), but vLLM extracts a larger DFlash speedup (2.90x vs our 2.47x) because its draft step is fully device-resident + CUDA-graphed (ours host-orchestrates 13 downloads/step) + slightly higher acceptance (~4.3 vs ~3.6 draft tokens/step). The DONE speed bar (ours >= vLLM) is NOT met; closing it = the device-resident draft rewrite + FULL CG (D6 part 2). STRICT-4/4 is bf16-IRREDUCIBLE β€” the draft KV cache is bf16 not fp8 (`torch_utils.py:398` `auto`β†’model dtype; the D0 "fp8-KV" was the backend name), the D3 golden already compares pre-storage bf16 (residual = sub-ULP kernel noise), and a fused multi-layer KV GEMM is per-element invariant to our per-layer GEMMs (bit-exact needs vLLM's exact kernels) β‡’ the ratified near-tie gate is the FINAL correctness form (no fused-KV code landed). FULL CG + persistent-paged-KV BLOCKED on a device-resident draft-path rewrite (the D5 path does 13 deviceβ†’host downloads/step) β€” the remaining throughput-parity increment. Inertness by construction (D5 binary; SACRED 235/235 + MTP 9/9 stand). Row STAYS `ACTIVE` (correctness-final + c1-speed-measured). Evidence tool `scripts/spec/vllm_dflash_timing.py`. **D7 `CLAIM-DFLASH-D7` 2026-07-27 (source-owning) β€” within-step draft forward DEVICE-RESIDENT:** `PrecomputeContextKVDevice` keeps per-layer K/V on device + `ForwardBlockLogitsWithContext` builds `[context;block]` with `vt::IndexCopy`/`IndexSelect` (removes ~30 Dβ†’H `Download`s/step), bit-identical (identity `bf16↔f32` round-trips replaced, no float op changed) β€” e2e `test_qwen27_dflash_spec_decode` 27/27 SAME tokens (2/4 STRICT + 2/4 near-tie at identical divergence points, acceptance 19/39/29/25), SACRED 235/235 + MTP 9/9, CUDA `-Werror` clean, compute-sanitizer 0 (198412; no new kernel). The direct old-vs-new A/B = **+2.0% output-tput (IN-NOISE)** β‡’ **D6's "downloads = the ~14% gap" REFUTED by measurement**; ours 19.68 tok/s STILL **~33% BELOW** vLLM-DFlash-ON 29.2 (reconstructed 8-prompt set; OFF parity our 9.97 β‰₯ vLLM 9.66); residual re-attributed to **acceptance** (ours 2.49 vs vLLM ~3.13 accepted draft-tok/step, bf16-irreducible) + per-step **context-KV RECOMPUTE** (O(contextΒ²); needs the cross-step persistent paged draft-KV store) + eager-vs-graphed. SPEED BAR NOT met; STAYS `ACTIVE` **D9 `CLAIM-DFLASH-D9` 2026-07-27 (source-owning) β€” PERSISTENT PAGED DRAFT-KV LANDED (bit-identical, +22.7%):** `qwen3_dflash.{h,cpp}` `AppendContextKVHost` (project ONLY newly-accepted rows β†’ per-layer bf16 K/V) + `ForwardBlockLogitsWithPrecomputedKV` (upload the persistent store, NO re-projection) share the core `ForwardWithCtxKVDev` with the old recompute; `runner.cpp::propose_drafts_dflash` swaps the O(contextΒ²) per-step recompute (`dflash_ctx_feats_`) for an append-only per-request `dflash_kv_store_` (rollback=don't-append). NO new CUDA kernel; config-gated. Bit-identity: CPU `test_dflash_propose` 2 new D9 cases = exact float equality vs full recompute; GPU e2e `test_qwen27_dflash_spec_decode` **27/27 SAME tokens** (acceptance 19/39/29/25, same divergences); SACRED 235/235 + MTP 9/9 byte-identical; CUDA `-Werror` clean. **A/B (c1, 8 prose+codeΓ—256 tok, `benchmark_binding=true`):** ours-ON **25.75 tok/s** (was 20.99, +22.7%) / 38.40 ms / acc 3.68/step vs vLLM-ON graphed **28.09** / 35.60 / acc 3.31 = **0.917Γ—** (was 0.69Γ—). **Part 1:** same-trajectory per-step acceptance == vLLM EXACTLY (ratio 1.00 on the 2 token-identical prompts) AND ours realized acceptance (3.68) > vLLM (3.31) β‡’ D8's "bf16 acceptance ceiling" REFUTED (confound). Residual ~8% = eager-vs-graphed ONLY; the FULL uniform-(1+k) CG (device paged-KV store + paged attn, new-CUDA) is the SOLE un-landed increment. **D12/D13 `CLAIM-DFLASH-D12/D13` 2026-07-27 β€” fixed-capacity paged draft-KV store (`vt::DFlashPagedBlockAttention`) + the draft-step CUDA graph landed (capture-correctness PROVEN replayed==eager bit-identical); c1 0.917Γ—β†’0.978Γ—.** **D14 `CLAIM-DFLASH-D14` 2026-07-27 β€” SPEED GATE MET β†’ `DONE`:** nsys attributed the D13 ~2% residual to the from-scratch `DFlashPagedBlockAttentionKernel` draft attention (1.8% of GPU time, ~460 us/call; the draft bf16 GEMMs run `cutlass_80_wmma` in BOTH engines, NOT the gap); ported it to a WARP-scoped `__shfl_xor` online-softmax variant (mirrors the shipped `AttentionWarpKernel`; `src/vt/cuda/cuda_ops.cu` `DFlashPagedBlockAttentionWarpKernel`, default ON, `VT_DFLASH_ATTN_BLOCK=1` = the bit-identical D12/D13 block kernel) β†’ draft attn 242.9β†’77.9 ms (3.1Γ—), our-ON c1 28.60β†’29.32 tok/s. FINAL same-session 3-rep A/B: our-ON 29.42/29.27/29.32 vs vLLM-ON 29.240/29.247/29.233 β€” our WORST > vLLM's BEST, 1.003Γ— β‡’ **β‰₯vLLM MET.** Correctness UNCHANGED (e2e 27/27 graph==eager, acceptance 19/39/29/25 identical, 1629 draft accepted identical warp-vs-block; CUDA==CPU 795648/795648 + compute-sanitizer 0), inertness SACRED 235/235 + MTP 9/9. Correctness-complete (ratified near-tie) AND at/above vLLM throughput. Anchors: [cuda_ops.cu](../src/vt/cuda/cuda_ops.cu#L1433) + [test_ops_dflash_paged_block_attn](../tests/vt/test_ops_dflash_paged_block_attn.cpp#L79) + [ledger](parity-ledger.md#L738); closing commit `164453a2` (claims CLAIM-DFLASH-D0D1..D14 recorded in coordination.md + ledger). | `489a7544` | | `MODEL-SPEC-deepseek-v4-dspark-deepseek-v4-for-causal-lm` | `DSparkDraftModel` | `registry.py:592`; `vllm/models/deepseek_v4/__init__.py::DSparkDeepseekV4ForCausalLM` | speculative draft / target-dependent | draft runner; acceptance/sampling; FusedMoE/grouped GEMM; MLA/latent KV; sliding-window attention; DSpark | ☐ required | `INVENTORIED` | The DeepSeek-V4 destination stays `INVENTORIED` and unimplemented, but the STRING now has two: **BEYOND-PIN** vllm#52197 (merged 2026-08-17 at `7075ddac`) makes `DSparkDraftModel` name a QWEN3 draft whenever the draft's `model_type` is `qwen3`, which is what `RadixArk/Qwen3.8-27B-DSpark` declares. `SPEC-DSPARK-QWEN3-ROUTING` ([#1193](https://github.com/mudler/vllm.cpp/issues/1193), [spec](specs/dspark-qwen3-routing.md)) routes that pair to the landed `Qwen3DSparkModel` lane in `SpeculativeConfig::ResolveDsparkArchitecture` and REFUSES this DeepSeek-V4 destination by name from `LoadedEngine::ResolveSpecConfig`, rather than mirroring upstream's silent rewrite into a stub (Β§7 R2). Evidence: [test_speculative_dspark.cpp](../tests/vllm/config/test_speculative_dspark.cpp), [test_dspark_draft_routing.cpp](../tests/vllm/entrypoints/test_dspark_draft_routing.cpp) | unassigned | | `MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model` | `DSparkV41DraftModel` (the V4.1 DSpark drafter; V4.1 ships NO MTP head, so this is its only speculative arm) | **NOT AT THE PIN.** Registered on vLLM `main` at `e77daef89e`: `model_executor/models/registry.py:647` `"DSparkV41DraftModel"` -> `vllm.models.deepseek_v4_1`. At our pin `e126687a9a` the name does not occur (`grep -c` = 0); the pin's own DSpark row is `DSparkDraftModel` at `registry.py:624`, a DIFFERENT class against the V4 target. `vllm/models/deepseek_v4_1/` contains no `mtp.py`, and there is no `DeepseekV41MTP` at any revision, so V4.1 replaces V4's MTP+DSpark pair with DSpark alone -- `config/speculative.py` branches on `deepseek_v41` to select this draft arch. | speculative draft / target-dependent | the `MODEL-MM-deepseek-v4-1-deepseek-v41-for-causal-lm` target first; DSpark block drafting (`dspark_block_size` 5, `dspark_target_layer_ids` [37,38,39], `dspark_markov_rank` 256, its own 128-expert MoE at top-3); spec-decode acceptance plumbing | [deepseek-v4-1-flash](specs/deepseek-v4-1-flash.md) | `BLOCKED` | **SCOPING ONLY, 2026-09-11: no product code.** Blocked on the same pin advance as its target, and additionally on the target existing at all -- a draft head is not portable before the model it drafts for. The `dspark_*` config keys are read from the release `config.json` @ `dba1be0a40`; the drafter's weights ship inside the 475.27 GiB release and inside the EXL3 checkpoint (where they stay MXFP4, not EXL3). | unassigned | +| `MODEL-DSV41-EXL3` | **A V4.1 CHECKPOINT-CAMPAIGN ROW, NOT A MODEL-SPEC TARGET.** The `bot-lab-21/DeepSeek-V4.1-Flash-EXL3-3.5bpw-Pollard` artifact @ `f129e31a81e1337aa33e129e2d847fc7e37c8733`, 428.49 GiB β€” the quant arm of the `MODEL-MM-deepseek-v4-1-deepseek-v41-for-causal-lm` target, and the V4.1 sibling of `MODEL-DSV4-EXL3`, the V4 campaign row | **THE HEADLINE FORMAT CLAIM DOES NOT DESCRIBE MOST OF THE FILE, verified by LFS oid rather than by the card.** Shards 1, 2 and 43-48 are BYTE-IDENTICAL to the deepseek-ai release β€” 199.88 GiB of it β€” and only 3-42 differ; `quantized_modules` covers the ROUTED EXPERTS ALONE, leaving fp8 dense and attention, MXFP4 shared experts and the DSpark drafter, fp8 engram tables and native fp4 KV as the release's own bytes. "EXL3 3.5bpw" is therefore an AVERAGE OVER PART OF ONE FILE, not a format for the file, and "Pollard" is a Hessian-aware bit-allocation RECIPE over per-expert widths: the bytes are turboderp EXL3 trellis, which is the `MODEL-DSV4-EXL3` format and not a new one | quantized arm of the V4.1 text/MoE backbone | the exllamav3 trellis kernels `MODEL-DSV4-EXL3` landed (`kExl3Gemm`, `kExl3HadR128`); the engram, MXFP8 32x32 UE8M0 linear, ratio-1/2 indexer, KV topology and DSpark families its non-trellis bytes need are OWED to W3a-W3f of the target row; and four MATCHED fleet devices, which this fleet does not have | [deepseek-v4-1-flash](specs/deepseek-v4-1-flash.md) (issue `ISSUE-LOCAL-01M29ATB6N2SFZD7JA6CCR2KXC`) | `BLOCKED` | **SCOPING ONLY, 2026-09-11: no product code, and none reachable from this artifact yet.** The tree's own answer to the target is a refusal: `src/vllm/model_executor/models/deepseek_v4_1_registry.cpp:109-135` throws "the safetensors weight loader is not ported" and names W8 as the wave that owes the loader, so NO V4.1 weights load in ANY format here and W1 makes the architecture RESOLVE only. Blocked on the vLLM-pin advance AND on hardware: 3.6x one GB10, the publisher's numbers are 4x DGX Spark TP4 and are unreproduced here, and the exllamav3 revision that produced the artifact is UNVERIFIED | unassigned | +| `MODEL-DSV41-GGUF-Q1_0` | **A V4.1 CHECKPOINT-CAMPAIGN ROW, NOT A MODEL-SPEC TARGET.** The `vcruz305/DeepSeek-V4.1-Flash-GGUF` `Q1_0` rung @ `543d86fd975044332d4366204eadd6da1b0cbeda`: 38.781 + 41.776 + 18.033 = 98.591 GiB against a 119 GiB unified pool, ~20.4 GiB clear. **It is the only published V4.1 artifact that fits one GB10, and the size is the only part of it that works** | `deepseek41` is in NO RELEASED ENGINE: absent from llama.cpp `b10451` and from `master`, `ggml-org/llama.cpp#28696` is OPEN, DRAFT and conversion-only, and the runtime is an out-of-tree branch. `vllm-gguf-plugin` has no DeepSeek adapter at all | the degenerate rung of the V4.1 GGUF conversions, and a READER CORRECTION rather than a port: a 1-bit block format keeps one scale per block and one sign bit per weight, the file stores `token_embd` ITSELF at ggml type 41, and the embedding table therefore returns plus-or-minus a single magnitude | no released engine that can load it, and no published rung that is both non-degenerate and fits; constructing one β€” keep `token_embd` and `output` out of the 1-bit format and spend the headroom on them β€” is a REQUANTIZATION question and is not scoped here | [deepseek-v4-1-flash](specs/deepseek-v4-1-flash.md) (issue `ISSUE-LOCAL-01M29ATQDVEQ7M2MP3VW79XQJP`; the format-level evidence correction rides on `QUANT-GGUF-Q1_0`) | `BLOCKED` | **SCOPING ONLY, 2026-09-11: no product code, and the target would refuse it in any case.** The producer root-caused the degeneracy as "the file, not the runtime" on `ggml-org/llama.cpp#28696`, measured 1.53 bits per weight over the backbone, and states Q2_K is the first rung with enough magnitude left β€” at 246.349 GiB, twice this box. Across the published rungs "fits" and "can say anything" DO NOT OVERLAP. This tree's own answer is `DeepseekV41GgufRefusal()`, thrown from `src/vllm/model_executor/models/deepseek_v4_1_registry.cpp:109-135` for any GGUF source, and the entrypoint's architecture dispatch throws the same string, so whichever door a GGUF arrives at the message is identical | unassigned | | `MODEL-SPEC-qwen3-dspark-qwen3-dspark-for-causal-lm` | `Qwen3DSparkModel` | `registry.py:593`; `vllm/model_executor/models/qwen3_dspark.py::Qwen3DSparkForCausalLM` | speculative draft / target-dependent | draft runner; acceptance/sampling; DSpark | ☐ required | `INVENTORIED` | none | unassigned | | `MODEL-SPEC-laguna-dflash-dflash-laguna-for-causal-lm` | `DFlashLagunaForCausalLM` (v0.25.0 target-pending) | v0.25.0 target `registry.py:598`; `vllm/model_executor/models/laguna_dflash.py::DFlashLagunaForCausalLM` @ `702f481` | speculative draft / Laguna targets | draft runner; acceptance/sampling; full/sliding attention; DFlash | ☐ required | `INVENTORIED` | none | unassigned | | `MODEL-SPEC-llama-eagle3-eagle3-llama-for-causal-lm` | `PEagleDraftModel`, `PeagleLlamaForCausalLM`, `Eagle3LlamaForCausalLM`, `Eagle3MiniMaxM2ForCausalLM`, `LlamaForCausalLMEagle3`, `Eagle3Qwen2_5vlForCausalLM`, `Eagle3Qwen3vlForCausalLM` | `registry.py:594-600`; `vllm/model_executor/models/llama_eagle3.py::Eagle3LlamaForCausalLM` | speculative draft / target-dependent | draft runner; acceptance/sampling; EAGLE/EAGLE3 | ☐ required | `INVENTORIED` | none | unassigned | diff --git a/.agents/specs/deepseek-v4-1-flash.md b/.agents/specs/deepseek-v4-1-flash.md index edcc235f1f..095566d354 100644 --- a/.agents/specs/deepseek-v4-1-flash.md +++ b/.agents/specs/deepseek-v4-1-flash.md @@ -10,9 +10,23 @@ Issues, one per row, all canonical-local: `ISSUE-LOCAL-01M29AV2VX9RBRF409F6YVX3MK` on `QUANT-GGUF-Q1_0` for the reader correction this scope uncovered. Base SHA: `a2532471b` -Matrices: [model-matrix.md](../model-matrix.md) (the two registry arches), -[kernel-matrix.md](../kernel-matrix.md) (the two checkpoint campaigns), +Matrices: [model-matrix.md](../model-matrix.md) (the two registry arches, the +DSpark draft row, and the two checkpoint campaigns below), [quantization-matrix.md](../quantization-matrix.md) (`QUANT-GGUF-Q1_0`). +**CORRECTED 2026-09-27: the two checkpoint campaigns are in `model-matrix.md`, +not `kernel-matrix.md`,** and this line used to say otherwise. Two rows decided +it. `MODEL-SPEC-deepseek-v4-1-dspark-v41-draft-model` β€” this spec's own +sibling row, same wave, same 2026-09-11 scoping pass β€” is declared in +`model-matrix.md`, as is `MODEL-DSV4-EXL3`, the V4 campaign row these two are +the successors of. `kernel-matrix.md` also pins its row count in two places +(`## Count invariants`: "exactly 52 practical kernel-family rows" and the +lifecycle tally, with `scripts/check-agent-record.py` holding the total), so +neither of these two rows can be added there without a second change to that +contract β€” which is a claim about cost, not a reason to file a model campaign +as a kernel family. Both rows were named by this header and by an open +canonical-local issue while appearing in NO matrix, so their issue records +failed `validate_issue_record` with "row is not canonical and claimable" and +`agent-issue-index.py` could not validate a single record. This spec landed as records only and carries the campaign's product code from W1 onwards; no row here is `READY`. The header names **five** rows. Of the four diff --git a/.agents/specs/gliner2.5-decide.md b/.agents/specs/gliner2.5-decide.md index c18598ed93..fd03473b6b 100644 --- a/.agents/specs/gliner2.5-decide.md +++ b/.agents/specs/gliner2.5-decide.md @@ -11,11 +11,25 @@ and `vllm_decide` ABI from kev/laya. CPU + GPU (CUDA). ## Now -`SPEC` β€” spec written, issue open -(ISSUE-LOCAL-01M3APZC6GKX9ME6AE6336D3VY). Implementation not started. The gap -is verified: no classification-head model exists behind `/v1/systemone` for -DeBERTa-based encoders. The DeBERTa v2 encoder from MODEL-GLINER25 is already -implemented and is the key reuse. +`SPIKE` β€” spec written, issue open +(ISSUE-LOCAL-01M3APZC6GKX9ME6AE6336D3VY), and the implementation is **in +tree**, which this section previously denied. The earlier revision of these five +lines read "Implementation not started" and was wrong at the time it was +written: the same tree carries +`src/vllm/model_executor/models/gliner25_decide_registry.cpp` and +`gliner25_decide_head.cpp`, a suite registered at `tests/CMakeLists.txt:785` +(`tests/vllm/models/test_gliner25_decide.cpp:1-181`), and a live C API path at +`src/capi/vllm_c.cpp:2106-2160` routing `arch == "SpanExtractor"` to +`Gliner25DecideInference` at ABI v29. The row was ALSO absent from every matrix +while this spec's own heading named it, so `canonical_rows()` could not see it +and its issue record failed validation; `MODEL-GLINER25-DECIDE` is now declared +in `model-matrix.md` under `MODEL-TOKCLS`, beside the state the tree supports. + +The gap this spec opened is still the gap: no classification-head model existed +behind `/v1/systemone` for DeBERTa-based encoders, and the DeBERTa v2 encoder +from MODEL-GLINER25 remains the key reuse. What has changed is that this row is +`SPIKE` rather than `SPEC` β€” code and a registered suite in tree, no measured +end-to-end decision β€” and that the claim to the contrary is withdrawn. ## Scope diff --git a/scripts/check-agent-record.py b/scripts/check-agent-record.py index 78332cabd6..33554081d3 100644 --- a/scripts/check-agent-record.py +++ b/scripts/check-agent-record.py @@ -47,7 +47,16 @@ # IndexTTS2 (s2-mel + talker), Moss-TTS (delay + realtime), Qwen3-TTS, # Voxtral-Realtime and dSpark-V4.1 rows. Bumped because a new row EXISTS, # never to make a transition pass. - "MODEL": (AGENTS / "model-matrix.md", 384), + # 387 since 2026-09-28: +3 rows that three specs and three open canonical + # records already named but no matrix declared, so `canonical_rows()` could + # not see them and each record failed `validate_issue_record` with "row is + # not canonical and claimable": `MODEL-GLINER25-DECIDE` (the classification + # head sibling of `MODEL-GLINER25`, named by the heading of + # specs/gliner2.5-decide.md), and `MODEL-DSV41-EXL3` + + # `MODEL-DSV41-GGUF-Q1_0` (the two V4.1 checkpoint-campaign rows named by + # the specs/deepseek-v4-1-flash.md header). Bumped because the rows EXIST, + # never to make a transition pass. + "MODEL": (AGENTS / "model-matrix.md", 387), # 82 since 2026-07-21: +`QUANT-NVFP4-CT-W4A16` (compressed-tensors NVFP4A16 / # W4A16 β€” NVFP4 weights with BF16 activations, distinct from the existing # `QUANT-NVFP4-CT-W4A4` and `QUANT-NVFP4-MO-W4A16` rows in both scheme From 2e2b3ec1456ef36ab463f28d1c7d5b99b8899e38 Mon Sep 17 00:00:00 2001 From: Yoav Date: Thu, 1 Oct 2026 14:58:49 -0400 Subject: [PATCH 2/2] fix(GATE-ISSUE-INDEX-TABLE-SHAPE): close the canonical record its landed fix resolves The fix commit 15d48b092 declared the three orphan rows, corrected the two specs, and wrote this record's ## Resolution section, but left the header at State: OPEN with Closed: -. The record contract rejects that combination once a record is closed (scripts/issue_records.py:536-551: CLOSED requires a real Closed date and Resolution evidence), and the stale header contradicted the closure the PR itself claims, which the second review flagged. This commit flips State to CLOSED and dates Updated and Closed to 2026-10-01; the Resolution section already carries the evidence and matches the landed tree (SPIKE 11, BLOCKED 7, Total 387, and one checklist entry per new row). Gates re-run on the new head: check-agent-record.py and check-model-checklist.py both exit 0 (MODEL=387, checklist matches the row states); check-commit-style and check-commit-trailers pass on the range 15d48b092..HEAD. ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff] --- .../ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md b/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md index f721f3aca9..386ea77777 100644 --- a/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md +++ b/.agents/issues/GATE-ISSUE-INDEX-TABLE-SHAPE/ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS.md @@ -1,14 +1,14 @@ ID: ISSUE-LOCAL-01M3JVCCKJ2NXT4TTSHNZT2GJS Title: Three rows named by specs and by open issues are declared in no matrix, so agent-issue-index.py cannot validate any record Row: GATE-ISSUE-INDEX-TABLE-SHAPE -State: OPEN +State: CLOSED Kind: bug GitHub: - Mirror: PENDING Availability: FULL Created: 2026-09-27 -Updated: 2026-09-27 -Closed: - +Updated: 2026-10-01 +Closed: 2026-10-01 ## Problem