Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/completed/issue-index.md
Original file line number Diff line number Diff line change
Expand Up @@ -895,7 +895,7 @@ rather than merged. `scripts/check-agent-record.py` gates both.
| [#2307](https://github.com/mudler/vllm.cpp/issues/2307) | `SPEC-DFLASH2` | **`VT_DFLASH_BOUNDS_DEVICE` was read from `src/` and documented nowhere, so `check-env-doc` was RED on `origin/main` and every branch cut from it inherited a red preflight.** Introduced by `21ef6f053` ([#2274](https://github.com/mudler/vllm.cpp/issues/2274), [#2304](https://github.com/mudler/vllm.cpp/pull/2304)), read at `src/vllm/model_executor/models/qwen3_dflash_internal.h:375`. Reproduced on a CLEAN `origin/main` checkout with no local changes, so it was not an in-flight artifact; it failed both `check-env-doc` and `test_check_env_doc` under `scripts/agent-preflight.sh`. FOUND while landing [#2302](https://github.com/mudler/vllm.cpp/issues/2302) and FIXED IN FLOW, per AGENTS.md's rule that filing does not defer the fix. **Documented in `docs/ENVIRONMENT.md` rather than allowlisted**, following the convention its own family sets -- `VT_DFLASH_PAGED`, `VT_DFLASH_GRAPH`, `VT_DFLASH_ATTN_BLOCK` and `VT_FA2_DFLASH_BLOCK` are all documented there, and the allowlist is for kernel-internal tuning switches. The distinction is load-bearing here rather than clerical: the switch adds two `Copy` + `Synchronize` round-trips onto a path whose entire purpose is to avoid a sync, so enabling it changes the timing of the very thing `SPEC-DFLASH2` measures -- a DIAGNOSTIC run, never a speed run, and a reader has to be told that | bug |
| [#2312](https://github.com/mudler/vllm.cpp/issues/2312) | `SPEC-DFLASH2` | **`check-env-doc` was RED on `main`: `21ef6f053` (#2274 / #2304) landed `VT_DFLASH_BOUNDS_DEVICE` documented in its index row and its code comment but NOT in `docs/ENVIRONMENT.md`.** A BASE failure rather than a branch one — every branch cut after that commit inherits a red `scripts/agent-preflight.sh`, cannot reach a green gate before push, and the red is charged to whichever unrelated change runs the gate next; found exactly that way while gating [#2309](https://github.com/mudler/vllm.cpp/issues/2309). Documented beside the other `VT_DFLASH_*` entries as user-facing rather than allowlisted as kernel-internal, because the readback is a `Download` that SYNCHRONIZES on a path deliberately kept sync-free, so it changes timing as well as checking. Fixed in the same flow, as the in-flow rule requires | bug |
| [#2275](https://github.com/mudler/vllm.cpp/issues/2275) | `MODEL-MM-QWEN4-EXP` | **`LoadStackedExperts` implements ONE of the three residencies `GgufLoadPolicy::Route` can return for a stacked expert tensor, and silently expands the other two to bf16.** `Route` answers `kKeepQuant`, `kKeepF16` or `kNvfp4Fp4` for `GgufTensorRole::kStackedExpertWeight`, and the f16 arm is genuinely reachable for a rank-3 tower because `KeepF16KDim` gives that role a K dim (`src/vllm/model_executor/model_loader/gguf_keep_quant.cpp:59-60`). `LoadStackedExperts` (`src/vllm/model_executor/models/qwen4_exp_weights.cpp:148-167`) branches on `kKeepQuant` alone; both other values fall off the end into `return ExpandBf16(g, name, {e, n, k}, /*nk=*/true)` at `:166`. The loader therefore materialises a residency the policy did not ask for, and nothing refuses. **Loud today, and still a row.** At the released 512 x 640 x 2560 geometry the expansion is ~240 GB across the stack, so on every device this project owns it aborts as an allocation rather than answering wrongly — which is the good case. Two reasons it is still worth a number: the abort NAMES THE WRONG THING (an operator reads out-of-memory, not "the loader ignored your quantization policy", so the diagnosis leads away from the cause), and the POLARITY IS WRONG EVEN WHEN IT FITS — on a reduced expert count, a test fixture or a future smaller checkpoint the fall-through succeeds and hands back a bf16 tower where the policy asked for `kKeepF16` or NVFP4, and AGENTS.md §"Inherit vLLM defaults" records that a token gate cannot see a dtype that is too wide. **What closes it:** refuse by name any residency `LoadStackedExperts` does not implement, naming the requested residency and the tensor, which is the pattern AGENTS.md already requires of an unimplemented arm; supporting `kKeepF16` or `kNvfp4Fp4` for stacked experts is a separate implementation with its own red-first gate, and the refusal comes first so the gap is visible instead of discovered as an allocation failure. Found by W5d-4 ([#2249](https://github.com/mudler/vllm.cpp/issues/2249) item 4) while writing the MoE weight adapter and deliberately NOT fixed in that flow: it belongs to the W5a loader, not to an adapter branch, and a residency refusal needs its own red-before test. Recorded under `## Owed` in [`specs/qwen4-exp-flash-next.md`](../specs/qwen4-exp-flash-next.md) beside the NVFP4 debt, under campaign issue [#1978](https://github.com/mudler/vllm.cpp/issues/1978) | bug |
| [#2317](https://github.com/mudler/vllm.cpp/issues/2317) | `ENG-RECORD-CONFLICT-SURFACES` | **GitHub does not apply the `merge=union` driver, so `.agents/issue-index.md` is a repo-wide lock and the PR it blocks gets ZERO check-runs rather than a red.** Measured 2026-08-29 at `origin/main` `df024dce4` while unblocking [#2303](https://github.com/mudler/vllm.cpp/pull/2303). The mechanism is now PROVEN rather than hypothesised, on one case's three real blobs with one variable: `git merge-tree --write-tree 9d672e3b3 df024dce4` exits 0 because it honours `.gitattributes:7`, while `git merge-file -p ours base theirs` over the identical inputs exits 1 with a conflict hunk because it does not, and GitHub's mergeability computation is on the `merge-file` side. Neither side of that case edits or deletes a row (ours appends 1, `#2300`; theirs appends 6, `#2223 #2230 #2286 #2274 #2309 #2312`), so both obey the append-only rule perfectly and conflict anyway. **19 of 23 open PRs touch this file.** GitHub has computed mergeability for 9 of the 19; **8 are unmergeable, 7 of the 8 conflict on the index, and for 6 of those the index is the ONLY conflicting path** (#2311 #2310 #2303 #2301 #2267 #2248; #2281 adds one spec file; #2168 alone is unmergeable for an unrelated surface, `kernel-matrix.md`). The consequence is SILENT, which is the part [#883](https://github.com/mudler/vllm.cpp/issues/883) does not carry: an unmergeable PR gets no check-run at all, because GitHub schedules `pull_request` workflows against `refs/pull/N/merge` and stops recomputing that ref once the merge fails. On #2303 head `9d672e3b3` the check-run `total_count` is **0**, while `refs/pull/2303/merge` still EXISTS and resolves to `7b84f5cb2`, frozen at parents `a4b333329` (a pre-conflict `main`) and `bfb4f87f3` (the FIRST of the branch's three commits), so a reader sees a merge ref plus two runs stuck at `queued`/`in_progress` on that stale head since 20:25Z and concludes the queue is slow. Nothing will ever arrive. #2248 is in the same state. This contradicts `AGENTS.md:74` ("carries `merge=union`, so two branches that each append a row merge without a conflict"), the same sentence in this file's own preamble, and the same sentence in `scripts/check-issue-index-append-only.py:4-5`; the preamble copy is frozen as `INDEX_PREAMBLE` at `scripts/check-agent-record.py:1973`, so correcting it is itself a gated change. Under `AGENTS.md` §Records the second admitted shape ("a genuinely append-only file that can union-merge") does not exist on this forge, and the surface degrades into the lock the same section forbids ("If N concurrent pull requests edit file F, that file is a lock") with N = 19. NOT FIXED and deliberately: moving the index to a per-row surface changes `check-agent-record.py` and `check-issue-index-append-only.py` semantics and needs its own spec, a red-before test and a fresh review. Three options are sketched in the issue and none is chosen: per-row files under a directory read by glob (the shape `owed_issues()` already uses); a derived-at-read-time index; or keeping the file and removing the SILENCE rather than the lock. Owned by `ENG-RECORD-CONFLICT-SURFACES` ([#364](https://github.com/mudler/vllm.cpp/issues/364)), whose spec `retire-shared-record-surfaces.md` was measured at `d928e2c3` before this file had its present shape and lists it under neither in-scope nor out-of-scope | bug |
| [#2317](https://github.com/mudler/vllm.cpp/issues/2317) | `ENG-RECORD-CONFLICT-SURFACES` | **GitHub does not apply the `merge=union` driver, so `.agents/issue-index.md` is a repo-wide lock and the PR it blocks gets ZERO check-runs rather than a red.** Measured 2026-08-29 at `origin/main` `df024dce4` while unblocking [#2303](https://github.com/mudler/vllm.cpp/pull/2303). The mechanism is now PROVEN rather than hypothesised, on one case's three real blobs with one variable: `git merge-tree --write-tree 9d672e3b3 df024dce4` exits 0 because it honours `.gitattributes:7`, while `git merge-file -p ours base theirs` over the identical inputs exits 1 with a conflict hunk because it does not, and GitHub's mergeability computation is on the `merge-file` side. Neither side of that case edits or deletes a row (ours appends 1, `#2300`; theirs appends 6, `#2223 #2230 #2286 #2274 #2309 #2312`), so both obey the append-only rule perfectly and conflict anyway. **19 of 23 open PRs touch this file.** GitHub has computed mergeability for 9 of the 19; **8 are unmergeable, 7 of the 8 conflict on the index, and for 6 of those the index is the ONLY conflicting path** (#2311 #2310 #2303 #2301 #2267 #2248; #2281 adds one spec file; #2168 alone is unmergeable for an unrelated surface, `kernel-matrix.md`). The consequence is SILENT, which is the part [#883](https://github.com/mudler/vllm.cpp/issues/883) does not carry: an unmergeable PR gets no check-run at all, because GitHub schedules `pull_request` workflows against `refs/pull/N/merge` and stops recomputing that ref once the merge fails. On #2303 head `9d672e3b3` the check-run `total_count` is **0**, while `refs/pull/2303/merge` still EXISTS and resolves to `7b84f5cb2`, frozen at parents `a4b333329` (a pre-conflict `main`) and `bfb4f87f3` (the FIRST of the branch's three commits), so a reader sees a merge ref plus two runs stuck at `queued`/`in_progress` on that stale head since 20:25Z and concludes the queue is slow. Nothing will ever arrive. #2248 is in the same state. This contradicts `AGENTS.md:74` ("carries `merge=union`, so two branches that each append a row merge without a conflict"), the same sentence in this file's own preamble, and the same sentence in `scripts/check-issue-index-append-only.py:4-5`; the checker of the day pinned the expected preamble as a constant, `INDEX_PREAMBLE`, so correcting it was itself a gated change; W6 (commit 7dc2ef1ea) retired the live index and that pin with it, and correcting the preamble is today an edit to the archive and nothing else. Under `AGENTS.md` §Records the second admitted shape ("a genuinely append-only file that can union-merge") does not exist on this forge, and the surface degrades into the lock the same section forbids ("If N concurrent pull requests edit file F, that file is a lock") with N = 19. NOT FIXED and deliberately: moving the index to a per-row surface changes `check-agent-record.py` and `check-issue-index-append-only.py` semantics and needs its own spec, a red-before test and a fresh review. Three options are sketched in the issue and none is chosen: per-row files under a directory read by glob (the shape `owed_issues()` already uses); a derived-at-read-time index; or keeping the file and removing the SILENCE rather than the lock. Owned by `ENG-RECORD-CONFLICT-SURFACES` ([#364](https://github.com/mudler/vllm.cpp/issues/364)), whose spec `retire-shared-record-surfaces.md` was measured at `d928e2c3` before this file had its present shape and lists it under neither in-scope nor out-of-scope | bug |
| [#883](https://github.com/mudler/vllm.cpp/issues/883) | `ENG-RECORD-CONFLICT-SURFACES` | **GitHub reports `CONFLICTING` on `.agents/issue-index.md` while local git merges it cleanly, so the union driver #846 armed does not stop forge conflicts.** Filed 2026 from the LTX-2.5 landing campaign against PR #880 and never indexed here; its row is appended now, in the same commit as [#2317](https://github.com/mudler/vllm.cpp/issues/2317), because the issue that first observed this class was itself untracked by the surface it is about, a row-key scan for it having returned zero. It measured both directions on one case, `git merge-tree --write-tree` rc=0 with zero CONFLICT lines and `git merge --no-commit --no-ff` rc=0 with the path reported modified rather than unmerged, and its operator consequence stands unchanged: a `CONFLICTING` verdict from the forge is not evidence of a conflict, so reproduce it with a local `git merge` before acting on it. It deliberately left the mechanism unestablished ("the leading hypothesis is that GitHub computes mergeability without applying `.gitattributes` merge drivers. I did not verify that") and proposed a two-throwaway-branch experiment to settle it; #2317 settles it instead with no throwaway PRs, by running `merge-tree` and `merge-file` over the same three blobs so attribute handling is the only variable. Related #364, #595, #846, #573 | bug |
| [#2324](https://github.com/mudler/vllm.cpp/issues/2324) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **W5b-1 — the `Glm5NextTextAttention` block and the `OwnedTensor` -> host f32 bridge.** Split out of [#2241](https://github.com/mudler/vllm.cpp/issues/2241), which stays OPEN for W5b-2, because the two halves answer to different oracles: the block and the bridge answer to `transformers` v5.16.1 (`modeling_glm5_next.py:1064-1257`, sha256 `2092bbb4efa2a8087b74f4a4da37635c503fe1df9ae73f1e6e8342af8b4b8e8b`) and the llama.cpp [#27752](https://github.com/ggml-org/llama.cpp/pull/27752) container and need no cache over them, while the decoder layer and the forward answer additionally to `MakeGlm5NextKVCache` and the `[T, hc_mult, hidden]` manifold. Three defects a fluent wrong port produces, each with its own discriminating case: (1) the converter SPLITS `kv_b_proj` and transposes only the K half, so K contracts over its first inner axis and V over its second — at the published geometry a swap is a shape error, so the gate also carries a SQUARE case where the untransposed reading is shape-valid and merely wrong, separating by 2.9469 over all 900 values; (2) CROSS-LAYER top-k sharing — a `shared` layer reuses the previous full layer's selection, and a layer that recomputes RUNS and emits plausible tokens, so the gate carries BOTH the correct output and a recomputing port's, both from the same oracle run, and asserts ours is the first (320 of 800 values differ, max separation 1.52); (3) the all-masked padded row is filled with `finfo.min` and NOT `-inf`, so its softmax is uniform and its output finite — the `-inf` mutation reds 49 of 160 assertions. **The bridge ANSWERS O22's open residency question**: decode ONE DSA layer at a time, 0.4654 GiB, never the tower, whose expanded form is 426.72 GiB against a ~119.63 GiB box; the 1 GiB per-tensor ceiling sits EXACTLY 4x above the largest legitimate tensor and EXACTLY 9x below the smallest expert bank, is checked from the SHAPE before any allocation, and cannot make O19 / [#2260](https://github.com/mudler/vllm.cpp/issues/2260)'s `MoeGateUpSwiGLUGroupedCuda` throw reachable because no overload takes an expert bank. RED captured first from the plausible wrong port (9/14 cases, 63/150 assertions); green 14/14 + 160 and 13/13 + 96; 22 of 24 negative mutations kill their gate, one is recorded as EQUIVALENT with its reason, and the other is the `BlockToFloat`-null arm no input can reach in this build, whose PREMISE gate is armed instead and proved by removing the Q8_0 decoder. **NOT REACHED from a production entry point** — the staged-slice disclosure is spec `## Owed` O25 and the wiring is W5b-2's | feature |
| [#2327](https://github.com/mudler/vllm.cpp/issues/2327) | `PERF-QWEN35-STAGE-WEIGHTS` | **Qwen3.5's dense decode weights were ATS-retagged HOST pages, and reading them from the GPU cost 22.4% of decode throughput on GB10 — staging them as true device copies takes the 27B from 0.82x vLLM to PARITY on the measured prompt.** The target decode step is weight-bandwidth-bound (~50 GB swept per forward against GB10's 273 GB/s roof, a ~184 ms floor; activations at 8 query rows are ~100 KB), so a placement penalty on the weights IS the step time. `ResidentWeight` (`qwen3_5.cpp:1141-1170`) handed every decode GEMM a HOST pointer tagged with a device wherever `host_memory_is_device_addressable()` answers true, while vLLM's parameters are built inside the torch CUDA allocator (`model_loader/base_loader.py:52-54`) and are real device memory. MEASURED on `dgx:gpu0`, one boot, one binary at `21ef6f053`, Qwen3.8-27B bf16 + DFlash2 k=7 at concurrency 1, `VT_QWEN35_ALIAS_HOST_WEIGHTS` the only variable, four warm repeats per leg, interleaved A,B,A,B,A: alias ON medians **11.677 / 11.693 / 11.690** (the third being the TERMINAL CONTROL, matching the opening arm, so the run is admissible and drift cannot masquerade as effect) against staged **14.288 / 14.337** — **+22.4%**, with vLLM on the same prompt at **14.305** and staged ours at **14.31**. This is the mechanism `laguna.cpp:130-132` already records and already shipped default-ON for two other models (Laguna to 1.03x vLLM, DeepSeek-V4 to 1.144x); Qwen3.5 never received it, and `docs/ENVIRONMENT.md:259` names the decode risk in the alias's own row and leaves it open. **It CANNOT be a blanket default flip**, because this file also serves `Qwen3.8-2.4T-A95B` and [#1299](https://github.com/mudler/vllm.cpp/issues/1299) measured that checkpoint exhausting a 119.631 GiB box precisely BECAUSE the CUDA arm paid for its weights twice — host bytes plus a device copy. So the policy asks the BOX rather than answering once for the file: `DeviceStagingFitsBudget` stages only while `VT_QWEN35_STAGE_MIN_FREE_FRAC` (default 0.55) of total device memory remains free AFTER the copy, which a 50 GiB model on a 119.6 GiB box satisfies and the 2.4T model — already past the floor when its first dense weight arrives — never does, preserving #1299's invariant exactly. `VT_QWEN35_ALIAS_HOST_WEIGHTS=1` pins the retag, `=0` forces staging, and an unanswerable `DeviceMemoryInfo` keeps today's behaviour rather than guessing, because an unknown budget is not a licence to double a model's residency. The budget arithmetic is extracted as a PURE function so it is gateable without a device (a fake `vt::Backend` would need every pure virtual stubbed and would gate less), and five cases pin it: a fitting model stages, the #1299 shape is refused, the floor is load-bearing, an unanswerable budget refuses, and a weight larger than free memory refuses. Red-first by mutation — deleting the reserve floor fails exactly the #1299 case and the floor case, `BUILD rc=0 / TEST rc=1`. One case was WRONG when first written and is recorded as such: it asserted the EXACT floor boundary, which gates the binary representation of 0.55 rather than the policy, and was replaced by clearly-above and clearly-below cases plus a floor-moves-the-answer case. The declarations sit at the END of the header deliberately: the first draft inserted them mid-file and shifted `Fp8Weight` from `:628` to `:658`, breaking the recorded anchor three records cite and reding `check-agent-record` at 29 stale against a baseline of 28 | bug |
Expand Down
Loading
Loading