diff --git a/.agents/completed/issue-index.md b/.agents/completed/issue-index.md index af9706936..511a9d5c3 100644 --- a/.agents/completed/issue-index.md +++ b/.agents/completed/issue-index.md @@ -863,6 +863,8 @@ rather than merged. `scripts/check-agent-record.py` gates both. | [#2178](https://github.com/mudler/vllm.cpp/issues/2178) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **No llama.cpp RELEASE defines `glm5next`, so register a scoped PR-pinned oracle — and the two candidate PRs turned out to be COMPETING implementations that disagree on the architecture string, not the text half and the vision half of one stack.** Registers [`llama-cpp-glm5next`](../oracles/llama-cpp-glm5next.md) at `ggml-org/llama.cpp` PR #27752, object `8a8d0bcc4d5fdf024c457526245bec4bc3a12adc`, on the `llama-cpp-qwen4exp` precedent. Re-measured 2026-08-28 in a fresh bare clone whose only remote is `ggml-org/llama.cpp`, from refs and objects and never from a working tree: `ls-remote` heads `8a8d0bcc...` (#27752) and `9370c82d...` (#27773) agree with `gh api .head.sha`; `fetch --depth 1` serves both; `merge-base --is-ancestor refs/heads/master` is **rc=1** for both against a `b10451` control at rc=0; `git grep -il 'glm5next\|glm5_next' b10451` is **rc=1** tree-wide against a `glm4_moe` control returning nine files, and the same grep at `master` `50f068fff` is rc=1 too; `conversion/glm5next.py` is 4714 B and `src/models/glm5next.cpp` 55716 B at the pin, against a `no-such-file.py` probe at rc=128. **#27752 registers `LLM_ARCH_GLM5NEXT -> "glm5next"` (`src/llama-arch.cpp:87`) and has no vision at all (`grep -il glm5 -- tools/` rc=1); #27773 registers `LLM_ARCH_GLM5_NEXT -> "glm5-next"` (`:152`) with its own text graph `src/models/glm5-next.cpp` plus `PROJECTOR_TYPE_GLM5V -> "glm5v"`.** The published `unsloth/GLM-5.3-Flash-GGUF` at revision `d425e572fb96` declares `general.architecture = glm5next` in its first shard's header, which is #27752's spelling and our own converter's, so pinning #27773 would give a denominator that refuses both artifacts by name — one file, not two. **O4 corrected** in [`glm5-next-flash.md`](../specs/glm5-next-flash.md): the RELEASE half holds, the "no llama.cpp oracle" half no longer does, and what stays owed is the floor itself plus a vision denominator. **W6's vision denominator is owed and #27773 would not discharge it even out of draft:** the staged `mmproj-BF16.gguf` declares `clip.projector_type = glm5next` and `grep -c '"glm5next"' -- tools/` is rc=1 at BOTH heads, so no revision of llama.cpp can load the published mmproj today. `gateable = no` with #2178 owing the measurement: nothing was built and nothing was run, and a build is not a run. The run half is REACHABLE for the first time — UD-Q2_K_XL (101.2535 GiB over four shards, summed on the `UD-Q2_K_XL/` prefix rather than a substring match, which also catches a 9,429,920-byte `Shard_Rewrite/` sibling that is not a shard) was staging to the NAS when this row was written; the oracle file carries the per-shard state at a named instant because a live count in an append-only row is a drift-lock. `mmproj-BF16.gguf` is complete at sha256 `513c9bfc55898998186543caefc01626fb28e378b92f391018e1c3dd6655b113` computed locally. **The contrast worth carrying:** the opposite case landed the same day in [#2194](https://github.com/mudler/vllm.cpp/issues/2194) — for `glm_moe_dsa` stock `b10451` ALREADY carries `LLM_ARCH_GLM_DSA -> "glm-dsa"` (`src/llama-arch.cpp:85`, case `:1051`, enumerator `src/llama-arch.h:90`, graph `src/models/glm-dsa.cpp`, converter `conversion/glm.py:274-276`), re-verified in this same clone at rc=0, so that row needed no scoped file at all. The test is one command, not a judgement: does the pinned RELEASE name the architecture? Records only: no product code, no pin advance on `llama-cpp`, no build, no GPU lease | feature | | [#2218](https://github.com/mudler/vllm.cpp/issues/2218) | `MODEL-MM-QWEN4-EXP` | **The `hc_norm` gamma polarity disagrees between the loader and the device op, and a layer loop wiring them together scales by ~0.** `LoadNormBf16(..., unshift=true)` at `qwen4_exp_weights.cpp:264` stores the RAW HuggingFace gamma, centred on 0, by inverting the GGUF converter's baked `+1`. `vt::Qwen4ExpGatedResidual` documents the OPPOSITE convention — "hc_norm_w is vLLM's parameterization, i.e. ALREADY `1 + w_hf` … This op never adds 1" — so a layer loop that hands the loader's tensor straight to that op applies a near-zero scale, and the result reads as a corrupt checkpoint rather than as a wiring bug. The contradiction is visible AT THE LOAD SITE: the comment at `qwen4_exp_weights.cpp:258-263` argues FOR the fold, elementwise-corroborated on three published artifacts, immediately above the line that strips it. Nothing is broken today because `Qwen4ExpTextModel::Forward` does not exist; the moment the layer loop lands it must fold `hc_norm`, `norm_key`, `norm_query` and `norm_conv` through `vllm::qwen4_exp::HcNormWeightFromHf` first. NOT repaired in W5b-5, which hit the same shape and got it right by accident of scope: the QSA block's norms take the raw gamma under `RmsNormArgs::gemma = true`, which mutations M9/M10/M11 red. Owned by `MODEL-MM-QWEN4-EXP` and listed under `## Owed` in [`specs/qwen4-exp-flash-next.md`](../specs/qwen4-exp-flash-next.md); the layer-loop wave is where it gets fixed and gated. | bug | | [#2213](https://github.com/mudler/vllm.cpp/issues/2213) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **NoPE MLA and the DSA k-pool indexer — the geometry every later wave waits on.** W3 of [#1998](https://github.com/mudler/vllm.cpp/issues/1998). Two things, and each one fails quietly. (1) `MlaBlockDims::Validate` required every dimension `> 0` while `Glm5NextTextConfig.validate_architecture` REQUIRES `qk_rope_head_dim == 0` ("Expecting NoPE for the DSA attention layers"), so the two validators were exact complements over one field and no value satisfied both — O11, pinned executably in `test_glm5_next_scaffold.cpp` and now discharged: 0 is the ABSENT rotary, `head_size()` collapses to `kv_lora_rank` (512, not 576), and the block's rope branches become NOT TAKEN rather than zero-width work. Kimi-Linear is the near miss and is untouched: it keeps `qk_rope_head_dim = 64` and skips only the rotation. (2) `Glm5NextTextIndexer` scores LEARNED POOLED candidates, not raw tokens — `index_kpool` consecutive valid tokens compressed by a per-channel 4-way softmax with an intra-pool position embedding, `index_topk // index_kpool` pools selected, expanded back to raw indices, and the ragged tail appended raw and UNSCORED at width `index_topk + index_kpool - 1` = 2051. `deepseek_v4_dsa.cpp` has no pooling stage at all, so reusing it selects the wrong candidate set and yields plausible indices either way. `index_kpool` is **4** on the published artifact and 16 in the config class. Landed `src/vllm/model_executor/models/glm5_next_dsa.{h,cpp}` gated against goldens RUN out of transformers v5.16.1 at seq_len 21 vs index_topk 8 — STRICTLY past the threshold, because at or below it a top-k selects everything and the pooling is unobservable — asserting SET equality of the selected indices over 17 discriminating rows with a smallest margin of 2.58e-3. SACRED inertness proven by the six-arm DeepSeek byte-identity probe, base `150b37852` vs head, all six fingerprints identical | feature | +| [#2223](https://github.com/mudler/vllm.cpp/issues/2223) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **The 288 routed + 1 shared expert MoE, and the heterogeneous KV-cache spec — the first piece of this row that a production entry point REACHES.** W5 of [#1998](https://github.com/mudler/vllm.cpp/issues/1998). Two deliverables. (1) `src/vllm/model_executor/models/glm5_next_moe.{h,cpp}` BINDS rather than reimplements: the router is `vt::MoeRouterTopK`'s grouped `noaux_tc` arm and the epilogue is `deepseek_v4::ClampedSwiGLU` at `alpha=1, beta=0`, which is `_apply_gate`'s "Simple swiglu instead of alpha" line for line. Gated at the PUBLISHED 288/top-8 against goldens RUN out of `Glm5NextTextTopkRouter.forward` at transformers `v5.16.1`, asserting SET equality of the selected experts and PRINTING the separation margin (smallest 1.84e-3 over four tokens) — top-k error is bimodal, so a tolerance passes a wrong selection whose values happen to be close. Five silent-failure axes are each a killed mutation: sigmoid vs softmax scoring, `routed_scaling_factor` dropped, `norm_topk_prob` dropped, the `e_score_correction_bias` dropped (the bias SELECTS, the unbiased score WEIGHTS), and the factor applied TWICE by also passing it to `vt::MoeCombine`. (2) `MakeGlm5NextKVCache` replaces a refusal with THREE published groups — an `MLAAttentionSpec` at **512** for the 11 DSA layers (NOT the 576 every DeepSeek variant and Kimi-Linear publish: `qk_rope_head_dim` is ZERO here and upstream requires it), ONE uniform `MambaSpec` for the 34 KDA layers whose conv state is `conv_kernel_dim` = **4** columns wide and NOT `K-1` (`cache_utils.py:1015-1024` allocates it that wide and `causal_conv1d_update` reads `state_len` back off it; `kimi_linear_registry.cpp:157` publishes `K-1` for ITS model and copying that across hands the runner a cache one column short), and an `MLAAttentionSpec` at **257** = `2*index_head_dim + 1` for the indexer side cache at `compress_ratio` **1**, because the k-pool compresses at READ time and not in the store — the opposite of `MODEL-MM-QWEN4-EXP`'s QSA side cache. **REACHED**: the cases enter through `ModelRegistry::Resolve` and the `make_kv_cache` factory hook; unwiring `.make_kv_cache` REDS the gate, and DELETING the row is a `-Werror=unused-function` build error, so the toolchain itself proves the factory is the only reference. Measured on the way, and SUPERSEDED before this wave landed: the production loader run against the staged `unsloth/GLM-5.3-Flash-GGUF` rev `d425e572f` UD-Q2_K_XL arm opened the file, resolved `glm5next`, walked the 4-way split and stopped on `blk.3.ffn_gate_exps.weight has unknown ggml type id 17` (IQ2_XS). That was true when W5 measured it on 2026-08-29 and is not true now: [#2245](https://github.com/mudler/vllm.cpp/issues/2245) landed the IQ2_XS and IQ4_XS decoders and W5c ([#2242](https://github.com/mudler/vllm.cpp/issues/2242)) resolves all 1383 backbone tensors of that artifact, so the reading is kept as the measurement it was rather than as a live claim. The decoder layer, the DSA attention block and the assembled text forward are NOT in this wave and are carried as O23 | feature | +| [#2230](https://github.com/mudler/vllm.cpp/issues/2230) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **Three refusal messages named LANDED waves as owing, and one denied an artifact that exists — and the gate was PINNING all three.** Fixed IN FLOW under W5 of [#2223](https://github.com/mudler/vllm.cpp/issues/2223). (1) The forward refusal read "W3 the NoPE MLA block -- `MlaBlockDims::Validate` still refuses `qk_rope_head_dim == 0`", which W3 (#2213, `e511a614b`) made false by relaxing exactly that validator; it named W2's sigmoid forget gate and W4's unweighted mHC collapse as owed too, both landed (`199c44578`, `6c715de00`). W1 wrote the message and no later wave touched the file — `git log --oneline -- src/vllm/model_executor/models/glm5_next_registry.cpp` ends at W1's `47a2b35a5`. (2) The GGUF loader refusal read "NO `.gguf` of this model exists anywhere ... (O7)"; `unsloth/GLM-5.3-Flash-GGUF` rev `d425e572f` is published and four arms are staged. (3) The KV-cache refusal said the KDA layers carry "three separate conv states"; they carry ONE — the checkpoint's `self_attn.{q,k,v}_conv1d` concatenate into one grouped depthwise conv (`modeling_glm5_next.py:620-628`, `glm5_next_kda.h` "THREE LAYOUT FACTS"), and a spec written from that sentence would triple the group. THE MECHANISM: `test_glm5_next_scaffold.cpp` asserted all three sentences, so the gate passed *because* nothing had corrected them — a refusal message is this row's only user-visible surface and its assertions were pinning stale text rather than checking it. The repair adds the NEGATIVES (`MlaBlockDims::Validate still refuses` and `NO `.gguf` of this model exists` must NOT appear) so a revision that reintroduces either reds | bug | | [#2234](https://github.com/mudler/vllm.cpp/issues/2234) | `SPEC-DFLASH2` | **The batched-lane spec's `## Now` told a reader that L2 must NOT be merged, and L2 had been on `main` since `150b37852`.** `scripts/now.py` renders a row's live position from `## Now`, so the derived surface reported a landed change (#2212) as an unmergeable branch — the same defect class as [#2199](https://github.com/mudler/vllm.cpp/issues/2199), where a section written before a wave landed was never reconciled by the landing. Record-only repair, no product code. Three further claims had drifted and are marked DISCHARGED in place rather than deleted, so a later reader can tell "done" from "never written": the seam policy item landed as `c9b2049bc` (#2207), which is what makes a quantized gate-up arm reachable for the draft at all and so is a precondition of [#2224](https://github.com/mudler/vllm.cpp/issues/2224); O3 was already closed in `dflash2-batch-propose.md:348`; and the stale-anchor bullet cited the `P == 1` gate as `:1614` when it is `:1716`, so the CORRECTION had drifted twice as far as the `:1577` it was written to fix, which is the argument for `.agents/porting.md`'s name-the-symbol rule stated twice over. `## Now` now records L2's measured **-11% on `fwd`** (35.19 -> 31.3 ms, terminal control matching to 1.1%) and states **L3, the batched capture lane, as the row's next gate**: at `P > 1` the draft forward is not capture-targeted, so at c=8 the term that is 76% of the draft phase runs EAGER, while vLLM replays a FULL draft graph at every batch size and pads to `max_num_reqs` with `PAD_SLOT_ID` — verified at the parity pin `5559679229`, `spec_decode/dflash/speculator.py:456-458` (`run_fullgraph`) and `:589` ("Pad per-request buffers to max_num_reqs for CUDA graph safety"). A porting gap under "mirror vLLM", not a new design. It also records that the binaries carry no tree identity — `vllm_version()` returns `0.0.3+cuda` for every commit because `VLLM_CPP_BUILD_VERSION` defaults to `PROJECT_VERSION` — so L2's build is identified by its KERNEL SIGNATURE instead (`DFlashAttnMmaKernel` mangling to `...fbll`, 13 params carrying `tiles_per_req`, against `...fbl` in the pre-L2 `build23`), which proves the feature is compiled in rather than that a directory was named after a SHA | bug | | [#2240](https://github.com/mudler/vllm.cpp/issues/2240) | `QUANT-GGUF-IQ2_XS` | **IQ2_XS (17) and IQ4_XS (23) — the last two GGUF dequantizers the staged GLM-5.3-Flash artifact needed, and the two the loader stopped dead on.** "UD-Q2_K_XL" names a target average, not a format: of that artifact's 1412 tensors only TWO are Q2_K, while 82 are IQ2_XS (the `ffn_gate_exps`/`ffn_up_exps` routed experts) and 3 are IQ4_XS, so `LoadedEngine::FromModelDir` refused at `blk.3.ffn_gate_exps.weight has unknown ggml type id 17` before any dequant code ran — the reader had no block stride for 17, and the switch had no decoder for either. Both ported 1:1 from llama.cpp `b10451` (`ggml/src/ggml-quants.c:2516` `dequantize_row_iq2_xs`, `:2743` `dequantize_row_iq4_xs`, `ggml/src/ggml-common.h:627` `iq2xs_grid`) and gated BYTE-FOR-BYTE against the oracle's own decoders over REAL bytes read out of the two tensors that failed. IQ2_XS is the middle member of a family of three same-shaped codebooks — 256 / 512 / 1024 entries — where reaching for the wrong table still runs and still produces plausible magnitudes, so the 512-entry grid carries an FNV-1a seal as well. IQ4_XS reuses `kValuesIq4nl` unchanged; its delta is the super-block scale layout, a 6-bit `ls` spliced from a `scales_l` nibble and a `scales_h` bit pair and then biased by -32. Also carries the record correction the issue asked for: `.agents/specs/glm5-next-flash.md` O5/O8 are about the converter's WRITE side and were read as meaning the i-quant lane was absent entirely. Owning row `QUANT-GGUF-IQ2_XS` (and `QUANT-GGUF-IQ4_XS`); found by W5 of [#1998](https://github.com/mudler/vllm.cpp/issues/1998) via [#2223](https://github.com/mudler/vllm.cpp/issues/2223) | feature | | [#2243](https://github.com/mudler/vllm.cpp/issues/2243) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **`glm5next.attention.head_count_kv` is a per-layer `array[i32]` in the published artifact and `Glm5NextHfConfigFromGguf` reads it as a scalar.** Found while landing [#2240](https://github.com/mudler/vllm.cpp/issues/2240): with IQ2_XS and IQ4_XS decoded, the production loader gets past the type-17 refusal, opens all four shards, sizes all 1412 tensors, and stops instead at `glm5_next gguf: key glm5next.attention.head_count_kv is not an integer`. The artifact stores the layer schedule there — length 46, `0` on the 35 KDA layers and `1` on the 11 DSA/MLA layers — and `swiglu_clamp_exp`/`swiglu_clamp_shexp` are per-layer `array[f32]` of the same length directly behind it. Filed rather than fixed in that flow because it belongs to this row's config/loader wave and not to a dequant change; listed under `## Owed` as O18 in [`specs/glm5-next-flash.md`](../specs/glm5-next-flash.md) | bug | @@ -880,7 +882,33 @@ rather than merged. `scripts/check-agent-record.py` gates both. | [#2277](https://github.com/mudler/vllm.cpp/issues/2277) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **The published GLM-5.3-Flash GGUF is `tokenizer.ggml.pre = "glm4"`, and our pre-tokenizer table refuses that name — this is where the loader stops once [#2268](https://github.com/mudler/vllm.cpp/issues/2268) is fixed.** Measured 2026-08-29 on one tree and one build directory, three legs of one probe object driven through `LoadedEngine::FromModelDir` on `device = kCPU` at `/mnt/nas_share/rc/ckpt/GLM-5.3-Flash-UD-Q2_K_XL/`, headers only: the baseline reader stops at `attention.key_length_mla - attention.key_length is -256 but rope.dimension_count is 0`; with the MLA convention fixed but `attention.linear_head_count` still required it stops at `missing metadata key glm5next.attention.linear_head_count`, one key along; with both fixed it stops at `tokenizer: unsupported tokenizer.ggml.pre "glm4"`, past config resolution entirely. `src/vllm/tokenizer/tokenizer.cpp::FromGguf` maps seven pre names — `qwen35`, `qwen2`, `llama-bpe`, the four GPT-4o names, `deepseek-llm`, the three DeepSeek-V3 names and `laguna` — and refuses the rest by name. `glm4` is what every GLM-4 / GLM-5 GGUF carries; shard 1's KV block states `tokenizer.ggml.model = gpt2`, `tokenizer.ggml.pre = glm4`, 154880 tokens and 321649 merges. **The splitting rule is free and the BOS is not.** llama.cpp maps `glm4` and `chatglm-bpe` to `LLAMA_VOCAB_PRE_TYPE_CHATGLM4` (`b10451:src/llama-vocab.cpp:2256-2258`), whose regex at `:398` is BYTE-IDENTICAL to `LLAMA_VOCAB_PRE_TYPE_LLAMA3`'s at `:289`, so `SplitPattern::kLlama3` is EXACT here rather than the "close approximation" that [#347](https://github.com/mudler/vllm.cpp/issues/347) and [#1924](https://github.com/mudler/vllm.cpp/issues/1924) each had to undo — compare the two byte strings in the fix rather than trusting this sentence. But the same branch sets `special_bos_id = LLAMA_TOKEN_NULL` (`:2259`) while the artifact states `tokenizer.ggml.bos_token_id = 154822`, so llama.cpp DISCARDS a BOS id the file carries; a port that reads it and prepends it emits one token no reference run emits, on every request, and a token gate built from our own tokenizer could not see it because both sides would agree. Scope: map both names onto the CHATGLM4 rule with the byte comparison recorded rather than asserted, mirror the `special_bos_id` suppression with a case that fails if a BOS is prepended, and gate through `FromModelDir` on a `pre = "glm4"` fixture so the refusal that moves is the production one. Recorded as O20 in [`specs/glm5-next-flash.md`](../specs/glm5-next-flash.md), which carries the paired measurement | bug | | [#2279](https://github.com/mudler/vllm.cpp/issues/2279) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **`FromGguf` never reads `tokenizer.ggml.add_bos_token`, so a GGUF that asks for a leading BOS silently gets none.** Found while adding the `glm4` pre name for [#2277](https://github.com/mudler/vllm.cpp/issues/2277) and deliberately not fixed in that flow: #2277's scope is one pre name, this is a property of every GGUF tokenizer this tree loads. llama.cpp reads the key at `b10451:src/llama-vocab.cpp:2585-2586`, and `add_bos` is the ONLY thing that decides the prepend (`:3382-3384`, `if (add_special && add_bos)`); `tok::Tokenizer::FromGguf` reads `tokenizer.ggml.bos_token_id` and stops there, leaving `template_bos_` at -1 so `EncodeWithSpecialTokens` reduces to `Encode` for every GGUF. Nothing is red today because no artifact this tree gates on states the key -- the staged `unsloth/GLM-5.3-Flash-GGUF` UD-Q2_K_XL carries 72 KV entries and it is not among them, parsed 2026-08-29 from shard 1's own KV block, so llama.cpp's `add_bos` stays at its `:1815` default `false` and our silence is the right answer there. It is already live in the other direction on the `llama-bpe` family, whose arm at `:2157-2159` sets `add_bos = true` where the `glm4` arm at `:2256-2259` sets nothing, masked only because that path has never been token-gated against llama.cpp with `add_special = true`. No gate can see this class of defect: a prompt short by exactly one leading token still decodes to fluent text, still has a valid shape, still loads and still generates, and a token gate built from our own tokenizer compares us against us. Scope: read `add_bos_token` (and `add_eos_token`, the same upstream block) defaulting to llama.cpp's `false`; decide what represents it, since `template_bos_` has the right meaning and the wrong provenance comment; a case per arm proving exactly one BOS when true, none when false or absent, both round-tripping; and enumerate which committed fixtures and staged artifacts declare the key so the blast radius is measured rather than assumed. Recorded as O21 in [`specs/glm5-next-flash.md`](../specs/glm5-next-flash.md) | bug | | [#2273](https://github.com/mudler/vllm.cpp/issues/2273) | `BACKEND-TENSTORRENT-QWEN35` | **The wall is the per-CQ-operation tt-metal stack, charged once per staging write, and a decode step pays it once per staged tensor.** W5 (#2244) made uploads allocation-free and the wall honestly did not move (−0.14%, noise), and the trace split W4's hypothesis: `allocate_mesh_tensor_on_device_with_topology` is 0.02% of the AFTER profile and the write stacks are identical in both arms, so allocation was never the wall. What remains is the fixed per-op tax — `Threadpool::PollForWork` 14.29%, `MetalContext::instance` 11.14%, `memcpy` 6.23%, `Cluster::get_chip` 5.90%, `read_cq_host_ptr` 5.27% + sub-slices — multiplied by the layer fan-in. The lever (W4 lever 3, deferred there as optional, re-derived as owed): batch per-layer staging — pack a step's staged host rows into one contiguous host block and issue ONE mesh-CQ write per step or layer group, so the per-op tax divides by the fan-in. Invariant: staging stays bit-identical — the sacred golden pair 16/16 STRICT and the full TT suite green; this wave changes speed, never tokens. `StagingStats` gains route counters for the new path; capture-unsafe host-write refusals keep their semantics; f32-conversion arms keep their declared dtypes; a batched/arena layout must state its restage semantics explicitly (W5 review aliasing awareness — same-geometry restage aliases the persistent buffer), and the route must be production-reachable, not test-only. Evidence owed: same-method before/after profile on the P150 (identical leg, JIT-discard per arm, one lock hold) plus a fresh benchmark-record entry; the attribution shifts or the lever is named unreachable with the trace that proves it. The tt-metal-side residual (cached context handles, amortized CQ polling) stays recorded as the upstream-shaped alternative. Owned by `BACKEND-TENSTORRENT-QWEN35`, next wave after W5 (#2244, landed via #2258) | perf | +| [#2286](https://github.com/mudler/vllm.cpp/issues/2286) | `MODEL-DSV4-DSA-COMPOSE` | **The DeepSeek-V4 DSA composition had no owning row, and the forward's own refusal message said so** (`src/vllm/model_executor/models/deepseek_v4.cpp:~738`: "The DSA port itself is OWED and has no owning row"). SCOPED 2026-08-29 by [dsv4-dsa-compose.md](../specs/dsv4-dsa-compose.md), read at the parity pin `5559679229`. It is the blocker between a DeepSeek-V4-Flash artifact that LOADS ([#2186](https://github.com/mudler/vllm.cpp/issues/2186)/[#2283](https://github.com/mudler/vllm.cpp/issues/2283)) and one that RUNS. **The two kernel primitives already have rows** (`KERNEL-ATTN-DSA-SPARSE-INDEX`, `KERNEL-ATTN-DSA-COMPRESSOR`, both `SPIKE`); what had no owner is the ASSEMBLY into `AttentionBlock` -- three layer shapes selected by `compress_ratio` (`attention.py:454-533`), all ending at `forward_mqa` then `_o_proj`. **THREE FINDINGS THE SCOPING ADDS.** (1) The 3-way stream overlap is PERFORMANCE, not correctness: `attention_impl` dispatches with `enable=aux_streams is not None` and ROCm runs the same work sequentially, so a sequential first wave MIRRORS upstream -- stated so a later reader does not "restore" the overlap believing correctness depended on it. (2) `coff = 1 + (compress_ratio == 4)` is a per-token ROLE selected by offset within the gathering window -- `head_offset = (tokens >= COMPRESS_RATIO) * HEAD_SIZE`, emission at boundary tokens only, the state cache holding TWO head-sized rows per token, so a token in the overlap has a DIFFERENT role in each of the two windows containing it. That is the whole of what "never recoverable from the tensor alone" means, and our loader already materializes the width correctly ([#1970](https://github.com/mudler/vllm.cpp/issues/1970)), making this a FORWARD change rather than a loader one. (3) The compressor is two stages with the second boundary-gated, and its RoPE is exactly specified -- GPT-J style, `is_neox_style=False` (interleaved, NOT split-half), on the LAST `rope_head_dim` elements, at position `(positions // compress_ratio) * compress_ratio`. **HARD ORDERING:** consumes `KV-DSV4-MULTICACHE` ([#1925](https://github.com/mudler/vllm.cpp/issues/1925)) and W1 cannot start before that row's W3 hands the forward the cache. **NOT GATEABLE AT OR BELOW 512 TOKENS**, because the one arm that caches today forces indexer and compressor OFF and is exact only while `seq_len <= index_topk` (=512) -- so every gate on this row must exceed it. Also records that `config.json`'s `compress_ratios` has **46** entries `{0:5, 4:21, 128:20}` while older records say "43 layers"; 43 is the trellis shard count, and W1 reconciles which number each claim means before writing code. NOT FIXED IN FLOW and deliberately: `AGENTS.md` requires the spec first, and a capability of this size needs agreement on scope before implementation waves start | enhancement | +| [#2283](https://github.com/mudler/vllm.cpp/issues/2283) | `MODEL-DSV4-EXL3` | **The DeepSeek-V4 carried tower's BF16-sourced half is still widened to f32 (~2.62 GiB), and W1d's ~97.7 GiB projection has never been observed.** Filed 2026-08-29 because W1d ([#2186](https://github.com/mudler/vllm.cpp/issues/2186), landed `c9ad53fee`) CLOSED its issue while `.agents/specs/model-dsv4-exl3.md` `## Owed` still pointed two live entries at it -- a reader following either landed on a closed issue. **(1)** The 108.59 -> ~97.7 GiB figure is arithmetic on the measured 108.59 / 26.64 split, not a load anyone has watched complete; the last real measurement (2026-08-28, `dgx:gpu0`, worker `rc-worker-4b8lj`, tree `525d2b991`) REFUSED, and nothing has re-run since. It falls due as an `rc`-leased `dgx:gpu0` measurement against the staged 100 GB artifact, and a load that completes is still not a forward that runs (#1961, #1970, #1976). **(2)** The carried tower's other half -- norms, embeddings, router, `BF16` on disk, 2.621 GiB -> 5.24 GiB at f32 -- is untouched, and the same "Inherit vLLM defaults" argument applies verbatim. NOT folded into W1d deliberately: W1d's nine fields had three consuming functions and one device vtable entry, while this half is read by the sampler and lm_head paths too (the embedding is held twice on device, #1946), so it is a wave with its own gate. Build on what W1d left: `vllm::HostBf16`, the inlined `vllm::HostBf16ToF32` (out-of-line `vt::BF16ToF32` + no LTO would cost a call per element in the GEMV inner loop) with its exhaustive 65536-pattern agreement case, `Dot`'s bf16 overload, generic `MatVec`/`Gemm`/`expert_f32`/`GroupedOutputLora`, and a `DeepseekV4HostResidentBytes` that now reads each field's own `value_type` under a mutation-proven gate. Spec [model-dsv4-exl3.md](../specs/model-dsv4-exl3.md) `## Owed` | bug | | [#2249](https://github.com/mudler/vllm.cpp/issues/2249) | `MODEL-MM-QWEN4-EXP` | **The interleaved-mRoPE cos\|sin table builder was `static`, so the `qwen4_exp` QSA half of the layer loop could only have a SECOND copy of it.** Item 5 of five prerequisites measured while attempting the loop ([#2031](https://github.com/mudler/vllm.cpp/issues/2031)); W5d-2 closes that item only, and the other four stay open on this issue. `BuildMropeCosSinHost` sat at `src/vllm/model_executor/models/qwen3_5.cpp:9472` with internal linkage and all three of its uses inside that translation unit, and mRoPE is the arithmetic where a duplicate diverges in silence — a wrong axis still produces plausible tokens. **Fixed by `include/vllm/model_executor/models/qwen3_5_mrope.h`**, which takes the SIMPLER of the two shapes this row has already used: `RunGdnBlockPaged` (#2110) and `RunMoeBlock` needed a public WRAPPER because their signatures name `StepDevInputs`, a type qwen3_5.cpp declares privately, while this signature names only `std::vector`, `int64_t` and `vllm::HfConfig`, so the whole extraction is the `static` keyword plus a declaration. The definition does not move: `sed -n '9473,9514p'` of the base-SHA file and of the head both sha256 to `259b1b932cae0611...`. **A byte-identical body is not on its own a value guarantee**, because the keyword that changed is exactly the one deciding which definition a caller binds to, so `tests/vllm/models/test_qwen3_5_mrope.cpp` pins 152 f32 BIT PATTERNS over four cases — interleaved and chunked at the same config and positions, T == 1 at another rotary_dim and rope_theta, and a small-`t`-section case where the `pair <= 3 * sec[k]` boundary decides differently — against what the FILE-STATIC produced at base `94de63ff5`, captured by compiling its `sed`-extracted text in a standalone harness. Bitwise and not an epsilon: a pure host computation over `std::cos`/`std::pow` has no reduction-order freedom, so a tolerance would hide the only defect an extraction can introduce. 26 pre-existing qwen3_5/qwen4_exp suites are identical in exit status and in case and assertion counts before and after (base `94de63ff5` against branch head `c1ccbac19`, both of which predate this branch's merge of `main`; that merge brings W5b's `test_qwen4_exp_forward`, which makes the same glob match 27 on the merged head and is NOT part of this pair, having existed at neither end of it) — but **FOUR of the 26 measure nothing on a host without the checkpoints**, and only `test_qwen35_paged_engine` (rc 77) says so: `test_qwen35_gguf_spec_decode` (3 cases / **0 assertions**, `SKIP: set VLLM_MTP_GGUF_MODEL`), `test_qwen3_5_vl_e2e` and `test_qwen3_5_vl_video_e2e` (1 case / **0 assertions** each, `SKIP: Qwen3.6-27B checkpoint absent`) each exit 0 and print `Status: SUCCESS!`. Those last two are the STRICT token-exact e2e gates on `VLGenerateCoreGdn`, the driver core holding the call sites the reachability mutation deletes, so on such a host the reachability evidence rests ENTIRELY on `test_qwen3_5_moe_vision` (7 cases / 38 assertions, of which one case reds). **One equivalent mutant is recorded rather than hidden:** upstream's `<=` in `mrope.py:60-63` cannot be told from a `<` here, because the guard already requires `pair % 3 == 1` while `3 * sec[k]` is divisible by 3 — the boundary is unreachable, and the mutation that DOES red it shifts the bound instead. Reachability proven by deleting both production call sites, which reds `test_qwen3_5_moe_vision`'s `..._uses_MRoPE_positions_not_plain_1d`; the new suite stays green under that deletion and says so in its own comment, because a unit case measures the function and never that anything reaches it | feature | | [#2257](https://github.com/mudler/vllm.cpp/issues/2257) | `ENG-MM-QWEN36-VL-FORWARD` | **The four Qwen3.5/3.6 VL greedy drivers have no production caller: `ModelRegistry::Forward` cannot route an image or video request to any of them.** `Qwen3_5VLGenerateGreedy`, `Qwen3_5VLGenerateGreedyVideo`, `Qwen3_5MoeVLGenerateGreedy` and `Qwen3_5MoeVLGenerateGreedyVideo` are DEFINED at `src/vllm/model_executor/models/qwen3_5.cpp:9892,9915,9960,9974` and declared in `qwen3_5.h` / `qwen3_5_dense.h`; a grep for the four names over `src/ include/ examples/ tools/ benchmarks/` returns those four definitions and their six declaration lines and NOTHING else, so **every caller is in `tests/`**. The registered factories for `Qwen3_5ForConditionalGeneration` and `Qwen3_5MoeForConditionalGeneration` (`REGISTER_VLLM_MODEL`, `qwen3_5_dense.cpp:283`) route the forward to `ForwardQwen3_5Dense`, which takes a `ModelForwardInput` and carries no multimodal hook, and `ModelRegistry::Forward` additionally refuses a non-null `multi_kv` (`model_registry.cpp:428-440`) that this architecture's three cache groups make the runner set. The tree already states the same condition for the sibling 4B driver at `include/vllm/entrypoints/openai/chat_mm.h:266-267` — the M2c driver "runs it standalone, outside `ModelRegistry::Forward`". So M3-b image and M3d video are gated e2e and correct, and **no user arrives at them**, which by AGENTS.md `## Nothing lands dead` makes every change inside `VLGenerateCoreGdn` or below it reached by a test and by nothing else. FOUND, not caused, while landing W5d-2 of [#2249](https://github.com/mudler/vllm.cpp/issues/2249), which gave `BuildMropeCosSinHost` external linkage: that wave's `## Owed` entry in `.agents/specs/qwen4-exp-flash-next.md` has to name who owns the hop above its call sites, and nothing tracked this gap. The condition PREDATES the extraction and is unchanged by it in either direction. Owned by `ENG-MM-QWEN36-VL-FORWARD`, which owns `BuildMropeCosSinHost`, the shared `VLGenerateCoreGdn` and the two 27B dense drivers; the two MoE drivers additionally sit under `MODEL-MM-qwen3-5-qwen3-5-moe-for-conditional-generation` and [#891](https://github.com/mudler/vllm.cpp/issues/891) | bug | | [#2242](https://github.com/mudler/vllm.cpp/issues/2242) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **W5c — the weight tower and `load_weights`: `Glm5NextForConditionalGeneration` LOADS.** The GGUF arm of the registry's `load_weights` hook now returns a real `Glm5NextLoadedModel` built by `LoadGlm5NextFromGguf`, so this architecture has a `LoadedModel` for the first time and the loader's refusal is gone from product output. The tower covers every tensor group the architecture declares — the KDA layer with its three separate depthwise convs, the NoPE MLA with the two SPLIT absorbed halves, the DSA k-pool indexer, the flat mHC pair at `(2 + hc_mult) * hc_mult`, the 288 stacked routed experts plus one shared, and the dense MLP on the leading three layers — and refuses BY NAME on a missing tensor, a shape disagreement or a non-negative `ssm_a`. Gated against the REAL published artifact with no asset: `tests/vllm/models/glm5_next_gguf_manifest.inc` freezes the 1412-tensor header table of `unsloth/GLM-5.3-Flash-GGUF` UD-Q2_K_XL at revision `d425e572fb9686125831f476129e51cea34bc5b4`, and the name map is accounted against it in BOTH directions — 1383 enumerated, 0 missing, 0 unexplained, 29 MTP-block tensors deliberately dropped, 1383 + 29 = 1412. `blk.45` is NOT built as a decoder layer, asserted three ways because each alone is satisfiable by a wrong loader: no `blk.45.*` name is enumerated, the file demonstrably HAS one, and the loader positively COUNTS the 29 tensors it skipped. Driven at the staged artifact through the same chain `LoadedEngine::FromModelDir` uses, headers only: all four shards open, the config resolves to 45 layers / 34 KDA / 11 DSA / hc_mult 4 / kpool 4 / NoPE, and every one of the 1383 names resolves at 41 MB peak RSS with no payload byte read. The residency the load would take, predicted by `PeekRoute` over those same names: 736 tensors keep their blocks at 98.260 GiB, 647 expand to bf16 at 0.446 GiB. NOT DONE HERE and named in the spec: the forward (W5b, [#2241](https://github.com/mudler/vllm.cpp/issues/2241)), the KV-cache spec, the vision tower (W6), the MTP head (O2) and the safetensors arm, all five still refusing by name. Campaign [#1998](https://github.com/mudler/vllm.cpp/issues/1998), spec [`specs/glm5-next-flash.md`](../specs/glm5-next-flash.md) §W5c | feature | | [#2291](https://github.com/mudler/vllm.cpp/issues/2291) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **The W7a converter and the published artifact disagree on three tensors, and one of the three is a silent value transform.** Found while implementing W5c ([#2242](https://github.com/mudler/vllm.cpp/issues/2242)), whose own scope sentence assumed they agreed, and fixed in the same flow. Read at source from llama.cpp PR [#27752](https://github.com/ggml-org/llama.cpp/pull/27752) head `8a8d0bcc4d5fdf024c457526245bec4bc3a12adc` (`conversion/glm5next.py`, sha256 `bfacba27746096e7bb3ca4a2549c9026d3475e226c7f3edf230c37ffadc7b6b3`) plus the `DeepseekV2Model` it inherits, and confirmed against the staged UD-Q2_K_XL header table. (1) `.dt_bias` is RENAMED to `.dt_proj.bias` before the generic map runs, so the file carries `blk.N.ssm_dt.bias` and no bare `ssm_dt`. (2) `kv_b_proj` is SPLIT into `attn_k_b` and `attn_v_b` with the k half TRANSPOSED, so the file carries two tensors at DIFFERENT shapes — ne `[256, 512, 64]` and `[512, 256, 64]` — and no `attn_kv_b.weight`; because `qk_nope_head_dim == v_head_dim == 256`, a fixture at equal head dims cannot tell a correct split from a swapped one, so the gate asserts both SHAPES and the nearer-own-half property rather than sizes. (3) `ssm_a` holds `-exp(A_log)`, not `A_log` — the dangerous one, because the tensor is present, the shape is right and the values are plausible floats, so nothing structural fires: a loader that inverts gets NaN on every KDA decay, one that does not runs a sign-flipped forget gate and generates fluent wrong text, and no oracle for this model runs on any device this project reaches to tell the difference. Fixed in the converter, in the C++ name map (the split needs its own 1:1 table, since one HF name maps to two GGUF names and a dict cannot carry one key twice) and in the new loader, which refuses a non-negative `ssm_a` by name. `tests/scripts/test_convert_glm5_next_gguf.py` was RED on the tensor set before the converter moved | bug | +| [#2300](https://github.com/mudler/vllm.cpp/issues/2300) | `ENG-MM-INPUT-PIPELINE` | **The GPU runner never sets `ModelForwardInput.mm`, so a Qwen3-VL server throws on the first forward step of every request, text or image.** Measured at `e541be98a`. `ForwardQwen3VLForConditionalGeneration` opens with `VT_CHECK(input.mm.has_value(), ...)` at `src/vllm/model_executor/models/qwen3_vl_registry.cpp:127`, and that forward is what a loaded Qwen3-VL resolves to: `REGISTER_VLLM_MODEL(qwen3_vl, "Qwen3VLForConditionalGeneration", ...)` at `:203` binds `.forward` at `:184`. The field it demands is `std::optional mm = std::nullopt` (`include/vllm/model_executor/models/model_registry.h:446`), and the runner's designated initializer at `src/vllm/v1/worker/gpu/runner.cpp:2234` names 16 fields and NOT `.mm` before calling `ModelRegistry::Forward` at `:2340`. Over the whole 4443-line file a grep for `mm_features`, `MultiModalForwardInput` or `.mm = ` returns 0, and so does a grep for `mm`, `multimodal` or `MultiModal`; `include/vllm/v1/worker/gpu/input_batch.h:90` records the worker input batch as a subset with "mm_features / generator / lora / prompt_embeds / pooling DEFERRED", so the features the field would be built from never reach the worker at all. The three writers of `.mm` in the tree are single-sequence drivers (`qwen3_vl.cpp:638`, `gemma4_mm.cpp:240`, `muse_glimmer_mm.cpp:358`), none of them the runner. **The refusing shape is a per-model choice, not a tree-wide one:** `gemma4_registry.cpp:145` and `muse_glimmer_registry.cpp:113` both guard with `if (input.mm.has_value())` and both carry the sentence "nullopt on every text step => the text path below is byte-identical", so under those two a runner step with `mm` unset runs the text path while Qwen3-VL throws. **STATICALLY DERIVED and NOT RUN:** no binary was built and no server was started, because the filing unit touches no product code; every claim is a `file:line` read plus the two grep counts, re-derived at the base SHA, and a runtime confirmation needs real safetensors weights because `LoadQwen3VLForConditionalGeneration` (`qwen3_vl_registry.cpp:95`) refuses any other source. Distinct from [#1358](https://github.com/mudler/vllm.cpp/issues/1358), which is the same root cause with a different symptom (the tower is loaded on the production path and read by nothing, costing memory rather than every request), and from [#2257](https://github.com/mudler/vllm.cpp/issues/2257), which is the Qwen3.5 and Qwen3.6 VL drivers having no production caller at all (`ForwardQwen3_5Dense` carries no multimodal hook, so it never reads `input.mm` and never throws). NOT fixed in flow: the repair is either a text arm in the registered forward or the runner building `mm` from staged encoder outputs, and the choice between them changes what an image request does, so it takes the surprising-fix path with its own spec, gate and independent review rather than an in-flow repair. Owned by `ENG-MM-INPUT-PIPELINE`, listed under `## Owed` in [`multimodal-track.md`](../specs/multimodal-track.md), and corrected on [`mm-serving.md`](../specs/mm-serving.md) | bug | +| [#2274](https://github.com/mudler/vllm.cpp/issues/2274) | `SPEC-DFLASH2` | **The DFlash2 paged draft route reads out of bounds EAGERLY at `max_num_seqs=1`, so the committed speed gate cannot measure our arm at all.** `scripts/dflash2-speed-gate.sh` on `main` refuses with `GATE_RC=2 / RESULT_PRESENT=no`: `vllm-cli` exits 134 with `vt cuda: cudaMemcpyAsync: an illegal memory access`, while the oracle arm completes in the same run. Reproducer needs no concurrency, no CUDA graph and no FA2 lane: `VT_DFLASH_PAGED=1 VT_DFLASH_GRAPH=0 vllm-cli --prompt "The capital of France is" --max-tokens 64 --repeat 5 --max-num-seqs 1`. **Repetition 1 completes 64 tokens and a LATER one dies**, so it needs state carried across requests — which is why 16-token single-shot probes survive. SEVEN candidates tested and excluded, each on an FA2-carrying build on one boot: the CUDA graph (`VT_DFLASH_GRAPH=0` still faults, so every earlier `cudaGraphLaunch` attribution was incidental), the FA2 block lane (`VT_FA2_DFLASH_BLOCK=0`), merged QKV (`VT_QWEN3_QKV_MERGE=0` plus an ON control on the same boot), **the whole seam adoption of [#2207](https://github.com/mudler/vllm.cpp/issues/2207) by building `c9b2049bc~1` = `f01fcc4c6`, which still faults and so exonerates it**, FA2 being compiled out (four earlier gate runs were measured on a binary with an EMPTY `CUDA FA2 compiled-arch manifest` because the staged lease script omitted `-DVLLM_CPP_CUTLASS_FETCH=ON` — the same defect that forced the 2026-08-24 retraction), the `max_seq_len` replay staleness fixed in `41dd3398a`, and the host-side bounds accounting (`ddd527f3f` added four shape-only checks that run on EVERY backend; they are silent on the failing configuration, so that class is excluded and the checks remain as a named refusal for whoever breaks the accounting later). `VT_DFLASH_PAGED=0` is the ONLY configuration that completes, and on it the gate PASSES: `GATE_RC=0`, **ours 12.361 tok/s vs vLLM 16.292, ratio 0.759 — 24% slower** at 0.789% SM-clock spread on one boot, which makes this issue NECESSARY BUT NOT SUFFICIENT for parity. `compute-sanitizer` cannot see the fault: the `vt cuda drop-in` layer initialises CUDA before the sanitizer interposes, so memcheck disables itself and reports THAT as its own "1 error" — two leases spent learning it, recorded so a third is not. The remaining lead is the one class the pre-existing guards cannot check on CUDA, because both are `kCPU`-guarded and their own comment names it: "the host values were right and the UPLOAD did not land on the tensor this call reads". An opt-in device readback (`VT_DFLASH_BOUNDS_DEVICE=1`, off by default because the read synchronizes and this call sits on the route's no-sync path) now asserts the DEVICE `seq_ext` and slot-map endpoints against the host derivation | bug | +| [#2309](https://github.com/mudler/vllm.cpp/issues/2309) | `ENGINE-HYBRID-PLACEMENT` | **The fp4-resident MoE refusal was lost when W3c moved every architecture onto the shared placement seam.** `RunMoeBlockPlaced` refused the arm; the refactor left that helper dead and the live `RunMoePlaced` path accepted it. Placing an fp4-resident arm uploads every expert at load and then computes on the host across the bus, so it is SLOWER than not placing — and **a token gate cannot see it**, because the tokens stay correct and only the placement is wrong. Refusal restored as a `placeable` / `unplaceable_reason` contract on the seam itself rather than in each caller, so a newly wired architecture inherits it; callers pass `layer.moe.expert_gate_fp4.empty()`. It fires only when a placement is in force (`placed_on != engine_device`), leaving an ordinary unplaced load untouched, since a guard that fired there would break every load. Proved by a COMPILING mutation: with the guard rewritten never to fire, `test_device_placement` goes red at 1 case / 2 assertions. The first mutation attempt failed to compile under `-Werror` on the unused parameters and the stale binary reported 19/19 SUCCESS, which is a passing mutant proving nothing — the mutant build's rc=0 is part of the evidence. Found while gating W3c, fixed in the same flow. Spec [`specs/hybrid-placement.md`](../specs/hybrid-placement.md) §W3d | bug | +| [#2302](https://github.com/mudler/vllm.cpp/issues/2302) | `MODEL-DSV4-DSA-COMPOSE` | **`dsv4-dsa-compose.md` said `KV-DSV4-MULTICACHE` W3 was OWED and that this row's W1 was blocked on it; W3 landed 2026-08-27 as `ca3dcda21` ([#2078](https://github.com/mudler/vllm.cpp/pull/2078), [#2068](https://github.com/mudler/vllm.cpp/issues/2068) CLOSED), and the real blocker is the ownerless W5.** Found 30 minutes after that spec merged as `a4b333329` ([#2287](https://github.com/mudler/vllm.cpp/pull/2287)), while verifying the row's readiness against the tree rather than its records. **The code names the wall itself** -- `ModelRegistry::Forward` (`src/vllm/model_executor/models/model_registry.cpp:430-440`) refuses with "no registered forward consumes a cache set keyed by layer name ... row KV-DSV4-MULTICACHE W5 owns the consuming forward" -- so a DeepSeek-V4 engine today constructs, publishes AND allocates all 167 buffers and refuses at the first forward. **This makes the ordering HARDER, not softer:** W3 had an owner and landed, while W4-W7 are proposals with no owner at all, so nothing in `MODEL-DSV4-DSA-COMPOSE` can begin until W5 acquires one. **Cause, which is the reusable part: two stale records agreed with each other and neither was the tree.** #1925's index row predates W3 and still describes the runner dropping groups silently; `kv-dsv4-multicache.md` `## Now` opened with "W3 (#2068) is claimed" while its OWN closing paragraph already said the engine allocates all 167 buffers and refuses naming W5. AGENTS.md `## History is git` says "Before you conclude anything about past work, read the spec and run `git log -S`", and `git log --oneline --grep '2068'` shows `ca3dcda21` immediately -- it was not run. FIXED by correcting the `## Dependencies` table, `## Now`, the `## Work breakdown` prerequisite and the `## Stop conditions` entry in [dsv4-dsa-compose.md](../specs/dsv4-dsa-compose.md), the `MODEL-DSV4-DSA-COMPOSE` row in [kernel-matrix.md](../kernel-matrix.md), and the misleading opening sentence in [kv-dsv4-multicache.md](../specs/kv-dsv4-multicache.md) so the next reader is not caught the same way. The [#2286](https://github.com/mudler/vllm.cpp/issues/2286) index row carries the original wrong claim and is NOT edited, because the index is append-only by policy; this row and the corrected spec are the record | bug | +| [#2307](https://github.com/mudler/vllm.cpp/issues/2307) | `SPEC-DFLASH2` | **`VT_DFLASH_BOUNDS_DEVICE` was read from `src/` and documented nowhere, so `check-env-doc` was RED on `origin/main` and every branch cut from it inherited a red preflight.** Introduced by `21ef6f053` ([#2274](https://github.com/mudler/vllm.cpp/issues/2274), [#2304](https://github.com/mudler/vllm.cpp/pull/2304)), read at `src/vllm/model_executor/models/qwen3_dflash_internal.h:375`. Reproduced on a CLEAN `origin/main` checkout with no local changes, so it was not an in-flight artifact; it failed both `check-env-doc` and `test_check_env_doc` under `scripts/agent-preflight.sh`. FOUND while landing [#2302](https://github.com/mudler/vllm.cpp/issues/2302) and FIXED IN FLOW, per AGENTS.md's rule that filing does not defer the fix. **Documented in `docs/ENVIRONMENT.md` rather than allowlisted**, following the convention its own family sets -- `VT_DFLASH_PAGED`, `VT_DFLASH_GRAPH`, `VT_DFLASH_ATTN_BLOCK` and `VT_FA2_DFLASH_BLOCK` are all documented there, and the allowlist is for kernel-internal tuning switches. The distinction is load-bearing here rather than clerical: the switch adds two `Copy` + `Synchronize` round-trips onto a path whose entire purpose is to avoid a sync, so enabling it changes the timing of the very thing `SPEC-DFLASH2` measures -- a DIAGNOSTIC run, never a speed run, and a reader has to be told that | bug | +| [#2312](https://github.com/mudler/vllm.cpp/issues/2312) | `SPEC-DFLASH2` | **`check-env-doc` was RED on `main`: `21ef6f053` (#2274 / #2304) landed `VT_DFLASH_BOUNDS_DEVICE` documented in its index row and its code comment but NOT in `docs/ENVIRONMENT.md`.** A BASE failure rather than a branch one — every branch cut after that commit inherits a red `scripts/agent-preflight.sh`, cannot reach a green gate before push, and the red is charged to whichever unrelated change runs the gate next; found exactly that way while gating [#2309](https://github.com/mudler/vllm.cpp/issues/2309). Documented beside the other `VT_DFLASH_*` entries as user-facing rather than allowlisted as kernel-internal, because the readback is a `Download` that SYNCHRONIZES on a path deliberately kept sync-free, so it changes timing as well as checking. Fixed in the same flow, as the in-flow rule requires | bug | +| [#2275](https://github.com/mudler/vllm.cpp/issues/2275) | `MODEL-MM-QWEN4-EXP` | **`LoadStackedExperts` implements ONE of the three residencies `GgufLoadPolicy::Route` can return for a stacked expert tensor, and silently expands the other two to bf16.** `Route` answers `kKeepQuant`, `kKeepF16` or `kNvfp4Fp4` for `GgufTensorRole::kStackedExpertWeight`, and the f16 arm is genuinely reachable for a rank-3 tower because `KeepF16KDim` gives that role a K dim (`src/vllm/model_executor/model_loader/gguf_keep_quant.cpp:59-60`). `LoadStackedExperts` (`src/vllm/model_executor/models/qwen4_exp_weights.cpp:148-167`) branches on `kKeepQuant` alone; both other values fall off the end into `return ExpandBf16(g, name, {e, n, k}, /*nk=*/true)` at `:166`. The loader therefore materialises a residency the policy did not ask for, and nothing refuses. **Loud today, and still a row.** At the released 512 x 640 x 2560 geometry the expansion is ~240 GB across the stack, so on every device this project owns it aborts as an allocation rather than answering wrongly — which is the good case. Two reasons it is still worth a number: the abort NAMES THE WRONG THING (an operator reads out-of-memory, not "the loader ignored your quantization policy", so the diagnosis leads away from the cause), and the POLARITY IS WRONG EVEN WHEN IT FITS — on a reduced expert count, a test fixture or a future smaller checkpoint the fall-through succeeds and hands back a bf16 tower where the policy asked for `kKeepF16` or NVFP4, and AGENTS.md §"Inherit vLLM defaults" records that a token gate cannot see a dtype that is too wide. **What closes it:** refuse by name any residency `LoadStackedExperts` does not implement, naming the requested residency and the tensor, which is the pattern AGENTS.md already requires of an unimplemented arm; supporting `kKeepF16` or `kNvfp4Fp4` for stacked experts is a separate implementation with its own red-first gate, and the refusal comes first so the gap is visible instead of discovered as an allocation failure. Found by W5d-4 ([#2249](https://github.com/mudler/vllm.cpp/issues/2249) item 4) while writing the MoE weight adapter and deliberately NOT fixed in that flow: it belongs to the W5a loader, not to an adapter branch, and a residency refusal needs its own red-before test. Recorded under `## Owed` in [`specs/qwen4-exp-flash-next.md`](../specs/qwen4-exp-flash-next.md) beside the NVFP4 debt, under campaign issue [#1978](https://github.com/mudler/vllm.cpp/issues/1978) | bug | +| [#2317](https://github.com/mudler/vllm.cpp/issues/2317) | `ENG-RECORD-CONFLICT-SURFACES` | **GitHub does not apply the `merge=union` driver, so `.agents/issue-index.md` is a repo-wide lock and the PR it blocks gets ZERO check-runs rather than a red.** Measured 2026-08-29 at `origin/main` `df024dce4` while unblocking [#2303](https://github.com/mudler/vllm.cpp/pull/2303). The mechanism is now PROVEN rather than hypothesised, on one case's three real blobs with one variable: `git merge-tree --write-tree 9d672e3b3 df024dce4` exits 0 because it honours `.gitattributes:7`, while `git merge-file -p ours base theirs` over the identical inputs exits 1 with a conflict hunk because it does not, and GitHub's mergeability computation is on the `merge-file` side. Neither side of that case edits or deletes a row (ours appends 1, `#2300`; theirs appends 6, `#2223 #2230 #2286 #2274 #2309 #2312`), so both obey the append-only rule perfectly and conflict anyway. **19 of 23 open PRs touch this file.** GitHub has computed mergeability for 9 of the 19; **8 are unmergeable, 7 of the 8 conflict on the index, and for 6 of those the index is the ONLY conflicting path** (#2311 #2310 #2303 #2301 #2267 #2248; #2281 adds one spec file; #2168 alone is unmergeable for an unrelated surface, `kernel-matrix.md`). The consequence is SILENT, which is the part [#883](https://github.com/mudler/vllm.cpp/issues/883) does not carry: an unmergeable PR gets no check-run at all, because GitHub schedules `pull_request` workflows against `refs/pull/N/merge` and stops recomputing that ref once the merge fails. On #2303 head `9d672e3b3` the check-run `total_count` is **0**, while `refs/pull/2303/merge` still EXISTS and resolves to `7b84f5cb2`, frozen at parents `a4b333329` (a pre-conflict `main`) and `bfb4f87f3` (the FIRST of the branch's three commits), so a reader sees a merge ref plus two runs stuck at `queued`/`in_progress` on that stale head since 20:25Z and concludes the queue is slow. Nothing will ever arrive. #2248 is in the same state. This contradicts `AGENTS.md:74` ("carries `merge=union`, so two branches that each append a row merge without a conflict"), the same sentence in this file's own preamble, and the same sentence in `scripts/check-issue-index-append-only.py:4-5`; the preamble copy is frozen as `INDEX_PREAMBLE` at `scripts/check-agent-record.py:1973`, so correcting it is itself a gated change. Under `AGENTS.md` §Records the second admitted shape ("a genuinely append-only file that can union-merge") does not exist on this forge, and the surface degrades into the lock the same section forbids ("If N concurrent pull requests edit file F, that file is a lock") with N = 19. NOT FIXED and deliberately: moving the index to a per-row surface changes `check-agent-record.py` and `check-issue-index-append-only.py` semantics and needs its own spec, a red-before test and a fresh review. Three options are sketched in the issue and none is chosen: per-row files under a directory read by glob (the shape `owed_issues()` already uses); a derived-at-read-time index; or keeping the file and removing the SILENCE rather than the lock. Owned by `ENG-RECORD-CONFLICT-SURFACES` ([#364](https://github.com/mudler/vllm.cpp/issues/364)), whose spec `retire-shared-record-surfaces.md` was measured at `d928e2c3` before this file had its present shape and lists it under neither in-scope nor out-of-scope | bug | +| [#883](https://github.com/mudler/vllm.cpp/issues/883) | `ENG-RECORD-CONFLICT-SURFACES` | **GitHub reports `CONFLICTING` on `.agents/issue-index.md` while local git merges it cleanly, so the union driver #846 armed does not stop forge conflicts.** Filed 2026 from the LTX-2.5 landing campaign against PR #880 and never indexed here; its row is appended now, in the same commit as [#2317](https://github.com/mudler/vllm.cpp/issues/2317), because the issue that first observed this class was itself untracked by the surface it is about, a row-key scan for it having returned zero. It measured both directions on one case, `git merge-tree --write-tree` rc=0 with zero CONFLICT lines and `git merge --no-commit --no-ff` rc=0 with the path reported modified rather than unmerged, and its operator consequence stands unchanged: a `CONFLICTING` verdict from the forge is not evidence of a conflict, so reproduce it with a local `git merge` before acting on it. It deliberately left the mechanism unestablished ("the leading hypothesis is that GitHub computes mergeability without applying `.gitattributes` merge drivers. I did not verify that") and proposed a two-throwaway-branch experiment to settle it; #2317 settles it instead with no throwaway PRs, by running `merge-tree` and `merge-file` over the same three blobs so attribute handling is the only variable. Related #364, #595, #846, #573 | bug | +| [#2324](https://github.com/mudler/vllm.cpp/issues/2324) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **W5b-1 — the `Glm5NextTextAttention` block and the `OwnedTensor` -> host f32 bridge.** Split out of [#2241](https://github.com/mudler/vllm.cpp/issues/2241), which stays OPEN for W5b-2, because the two halves answer to different oracles: the block and the bridge answer to `transformers` v5.16.1 (`modeling_glm5_next.py:1064-1257`, sha256 `2092bbb4efa2a8087b74f4a4da37635c503fe1df9ae73f1e6e8342af8b4b8e8b`) and the llama.cpp [#27752](https://github.com/ggml-org/llama.cpp/pull/27752) container and need no cache over them, while the decoder layer and the forward answer additionally to `MakeGlm5NextKVCache` and the `[T, hc_mult, hidden]` manifold. Three defects a fluent wrong port produces, each with its own discriminating case: (1) the converter SPLITS `kv_b_proj` and transposes only the K half, so K contracts over its first inner axis and V over its second — at the published geometry a swap is a shape error, so the gate also carries a SQUARE case where the untransposed reading is shape-valid and merely wrong, separating by 2.9469 over all 900 values; (2) CROSS-LAYER top-k sharing — a `shared` layer reuses the previous full layer's selection, and a layer that recomputes RUNS and emits plausible tokens, so the gate carries BOTH the correct output and a recomputing port's, both from the same oracle run, and asserts ours is the first (320 of 800 values differ, max separation 1.52); (3) the all-masked padded row is filled with `finfo.min` and NOT `-inf`, so its softmax is uniform and its output finite — the `-inf` mutation reds 49 of 160 assertions. **The bridge ANSWERS O22's open residency question**: decode ONE DSA layer at a time, 0.4654 GiB, never the tower, whose expanded form is 426.72 GiB against a ~119.63 GiB box; the 1 GiB per-tensor ceiling sits EXACTLY 4x above the largest legitimate tensor and EXACTLY 9x below the smallest expert bank, is checked from the SHAPE before any allocation, and cannot make O19 / [#2260](https://github.com/mudler/vllm.cpp/issues/2260)'s `MoeGateUpSwiGLUGroupedCuda` throw reachable because no overload takes an expert bank. RED captured first from the plausible wrong port (9/14 cases, 63/150 assertions); green 14/14 + 160 and 13/13 + 96; 22 of 24 negative mutations kill their gate, one is recorded as EQUIVALENT with its reason, and the other is the `BlockToFloat`-null arm no input can reach in this build, whose PREMISE gate is armed instead and proved by removing the Q8_0 decoder. **NOT REACHED from a production entry point** — the staged-slice disclosure is spec `## Owed` O25 and the wiring is W5b-2's | feature | +| [#2327](https://github.com/mudler/vllm.cpp/issues/2327) | `PERF-QWEN35-STAGE-WEIGHTS` | **Qwen3.5's dense decode weights were ATS-retagged HOST pages, and reading them from the GPU cost 22.4% of decode throughput on GB10 — staging them as true device copies takes the 27B from 0.82x vLLM to PARITY on the measured prompt.** The target decode step is weight-bandwidth-bound (~50 GB swept per forward against GB10's 273 GB/s roof, a ~184 ms floor; activations at 8 query rows are ~100 KB), so a placement penalty on the weights IS the step time. `ResidentWeight` (`qwen3_5.cpp:1141-1170`) handed every decode GEMM a HOST pointer tagged with a device wherever `host_memory_is_device_addressable()` answers true, while vLLM's parameters are built inside the torch CUDA allocator (`model_loader/base_loader.py:52-54`) and are real device memory. MEASURED on `dgx:gpu0`, one boot, one binary at `21ef6f053`, Qwen3.8-27B bf16 + DFlash2 k=7 at concurrency 1, `VT_QWEN35_ALIAS_HOST_WEIGHTS` the only variable, four warm repeats per leg, interleaved A,B,A,B,A: alias ON medians **11.677 / 11.693 / 11.690** (the third being the TERMINAL CONTROL, matching the opening arm, so the run is admissible and drift cannot masquerade as effect) against staged **14.288 / 14.337** — **+22.4%**, with vLLM on the same prompt at **14.305** and staged ours at **14.31**. This is the mechanism `laguna.cpp:130-132` already records and already shipped default-ON for two other models (Laguna to 1.03x vLLM, DeepSeek-V4 to 1.144x); Qwen3.5 never received it, and `docs/ENVIRONMENT.md:259` names the decode risk in the alias's own row and leaves it open. **It CANNOT be a blanket default flip**, because this file also serves `Qwen3.8-2.4T-A95B` and [#1299](https://github.com/mudler/vllm.cpp/issues/1299) measured that checkpoint exhausting a 119.631 GiB box precisely BECAUSE the CUDA arm paid for its weights twice — host bytes plus a device copy. So the policy asks the BOX rather than answering once for the file: `DeviceStagingFitsBudget` stages only while `VT_QWEN35_STAGE_MIN_FREE_FRAC` (default 0.55) of total device memory remains free AFTER the copy, which a 50 GiB model on a 119.6 GiB box satisfies and the 2.4T model — already past the floor when its first dense weight arrives — never does, preserving #1299's invariant exactly. `VT_QWEN35_ALIAS_HOST_WEIGHTS=1` pins the retag, `=0` forces staging, and an unanswerable `DeviceMemoryInfo` keeps today's behaviour rather than guessing, because an unknown budget is not a licence to double a model's residency. The budget arithmetic is extracted as a PURE function so it is gateable without a device (a fake `vt::Backend` would need every pure virtual stubbed and would gate less), and five cases pin it: a fitting model stages, the #1299 shape is refused, the floor is load-bearing, an unanswerable budget refuses, and a weight larger than free memory refuses. Red-first by mutation — deleting the reserve floor fails exactly the #1299 case and the floor case, `BUILD rc=0 / TEST rc=1`. One case was WRONG when first written and is recorded as such: it asserted the EXACT floor boundary, which gates the binary representation of 0.55 rather than the policy, and was replaced by clearly-above and clearly-below cases plus a floor-moves-the-answer case. The declarations sit at the END of the header deliberately: the first draft inserted them mid-file and shifted `Fp8Weight` from `:628` to `:658`, breaking the recorded anchor three records cite and reding `check-agent-record` at 29 stale against a baseline of 28 | bug | +| [#2331](https://github.com/mudler/vllm.cpp/issues/2331) | `ENG-HYBRID-PLACEMENT` | **`main` did not build: `tests/vllm/model_executor/test_placed_moe_roundtrip.cpp` called `vllm::RunMoeBlockPlaced`, which `866075b2f` ([#2313](https://github.com/mudler/vllm.cpp/pull/2313)) deleted.** Verified at `origin/main` rather than on a branch -- the symbol was declared in ZERO files under `include/`/`src/`, the test called it 4 times, and the test was registered at `tests/CMakeLists.txt:1587`. Red on `build-test-cpu`, `build-newest-gcc` and `sanitize-cpu (thread)`, so **every branch cut from main inherited it and no open pull request could go green**, since CI builds the merge commit. #2313 was right to delete the helper -- its own message says the W3c refactor had already left it dead -- it just did not delete the test keeping it compiling. **FIXED by deleting the obsolete file, and the two cases died for DIFFERENT reasons, which is why neither was ported.** (1) The fp4-resident refusal is SUPERSEDED by a strictly stronger case `866075b2f` added in the same commit (`test_device_placement.cpp:395-440`), which reaches a REAL cross-device placement the CPU-only original could not. (2) The byte-for-byte round trip is NOT PORTABLE by construction: it worked by passing `kCPU` as the placement device explicitly, while `RunMoePlaced` reads the device from `ActiveMoePlacementPlan()` and short-circuits with `if (placed_on == engine_device) return body(engine, dh)` -- same-device placement is INERT by design. **The round-trip gate is therefore OWED, and was owed before this**: the deleted file's own header said "It does NOT prove the cross-device arm ... it is the gate W3b still owes". What the deletion removed is a file that LOOKED like coverage while exercising a helper production had stopped calling. Recorded under `## Owed` in [expert-stream-device-slots.md](../specs/expert-stream-device-slots.md) as a cross-device byte-for-byte run. Deliberately not "ported" to a same-device assertion, which would only restate `return body(engine, dh)`. Found while landing [#2302](https://github.com/mudler/vllm.cpp/issues/2302), whose docs-only pull request was red on jobs its diff cannot affect | bug | +| [#2241](https://github.com/mudler/vllm.cpp/issues/2241) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **W5b-2a — the decoder layer, the mHC stream threading, `Glm5NextTextModel::Forward` and the KV binding.** The four deliverables that answer to `transformers` v5.16.1 (`modeling_glm5_next.py:1259-1329` and `:1409-1494`, sha256 `2092bbb4efa2a8087b74f4a4da37635c503fe1df9ae73f1e6e8342af8b4b8e8b`) and to `MakeGlm5NextKVCache`. **The `[T, hc_mult, hidden]` manifold is what this exists to get right**: `:1477` expands the embedding to four residual streams and nothing collapses them until `hc_head` at `:1493`, and a port that threads `[T, hidden]` and collapses early RUNS — finite, right-shaped, fluent, every sublayer gate on this row still green, and no end-to-end token gate for this model exists on this fleet to catch it downstream. Gated three non-overlapping ways: the per-layer `[B, S, 4, H]` streams asserted ELEMENTWISE; the oracle's own minimum pairwise stream separation (**6.4703**) asserted so those are shown to be discriminating rather than four copies of one value; and an EARLY-COLLAPSE DECOY produced by the SAME oracle modules with the manifold collapsed to its mean and re-broadcast after every layer, which the gate asserts we differ from by the oracle's own measured **2.4032**. The fixture is a FIVE-layer mixed schedule carrying all four control-flow combinations `:1261-1272` selects between plus a `shared` DSA layer, at the PUBLISHED `hc_mult` of 4, `seq_len` 12 against `index_topk` 8, row 1 left-padded by three. **The KV binding stores the LATENT and not what the reference stores**: upstream caches the EXPANDED K/V at `:1175-1179` (32,768 values per token per layer) and `DsaCache` stores the 512-wide `k_pass` plus the 257-wide packed indexer row, which is exactly what `MakeGlm5NextKVCache`'s groups 0 and 2 publish and what a case now asserts against the production `make_kv_cache` hook on the published `config.json`; the spec decided that and said the equivalence was to be PROVED, and the proof is a case asserting `ExpandKv(a ++ b) == ExpandKv(a) ++ ExpandKv(b)` EXACTLY, which NoPE is what makes true. An 8-token prefill plus 4-token continuation reproduces the reference's own `DynamicCache` run and agrees BIT-EXACTLY with the 12-token one-shot tail. **RED first, and the red was in the ORACLE CONFIGURATION rather than the port**: 4/10 cases and 7/1647 assertions failed with layer 0 green and every DSA layer red by 2.7 to 8.1, bisected against oracle intermediates (mHC pre plus `input_layernorm` agreed to 4.8e-07, the attention did not), and the cause is that `Glm5NextPreTrainedModel` sets `_supports_sdpa = True` so a default config resolves `_attn_implementation` to `sdpa`, whose `build_attention_mask_from_topk` returns a BOOLEAN mask (`:1249-1250`) where the eager arm returns the additive `finfo.min` one (`:1252-1256`) — the two DISAGREE on a left-padded query row where every key is masked, SDPA emitting 0.0 and eager's uniform softmax the mean of the values, measured 0.0 against 0.509. The generator now pins eager, which is the arm W5b-1 gated and the only interface `:1227-1228` says a 3-D per-(query, key) mask can reach; no token gate could see this, because the rows that differ are padding. **Fourteen negative mutations, all fourteen killing their gate — after a fifteenth finding that is the one worth reading.** The mutation truncating the attention's key range under a filled cache SURVIVED at 1647/1647, because its output is all-NaN and `NaN > x` is FALSE for every x, so the running maximum in the test's `MaxGap` helper never left its initial zero and an ALL-NaN FORWARD READ AS A PERFECT MATCH on every gap assertion in the file; `MinStreamSeparation` and the cached-tail loop were blind the same way. All three now treat a non-finite value as an INFINITE gap and report the count separately, the mutation then reds 3 assertions, and the suite grew 1647 -> 1656. Green 10/10 + 1656, with the eight sibling glm5 suites unchanged and green. **NOT REACHED from a production entry point, and this wave's own scope said it would be** — `ForwardGlm5NextForConditionalGeneration` still refuses by name, so O15, O16, O17, O23 and O25 are NOT discharged: `.agents/reachability.md` is explicit that "an intermediate hop that is itself unreached does not carry", and what changed is that five separate dead ends became ONE gated assembly point. Spec `## Owed` O26 carries that in the strong form and **W5b-2b owns the wiring**, which is why #2241 stays OPEN. W5b-2b's two halves are now scoped from measurement rather than guess: the weight bridge for the KDA, MoE, dense-MLP and mHC arms, whose 42 sparse layers' routed experts are ~1,150 GiB in f32 against a ~119.63 GiB box so an on-demand per-expert decode is the only shape that fits (`kBridgeTensorF32ByteCeiling` correctly refuses a 9.0 GiB bank today, which is O25's gate working as designed); and the engine binding, which is the SMALLER half and has a house pattern — `NemotronHForCausalLM` (`nemotron_h_registry.cpp:200-213`) and `KimiLinearForCausalLM` (`kimi_linear_forward.cpp:462-476`) both carry a host arm that ignores the paged caches and re-runs the whole prefix, and a survey of every `: public LoadedModel` found NO model keeping per-request state on it | feature | +| [#2329](https://github.com/mudler/vllm.cpp/issues/2329) | `PERF-QWEN35-STAGE-WEIGHTS` | **`VT_QWEN35_STAGE_MIN_FREE_FRAC` was read from `src/` and documented nowhere, so `check-env-doc` was RED on `origin/main` itself and every branch cut after `207c12932` inherited a failing preflight.** Introduced by `207c12932` ([#2327](https://github.com/mudler/vllm.cpp/issues/2327), [#2328](https://github.com/mudler/vllm.cpp/pull/2328)), read at `src/vllm/model_executor/models/qwen3_5_weights.cpp:218` and explained only in a code comment. Reproduced on `origin/main`'s own bytes, extracted with `git archive` into a clean directory with no branch involved: `check-env-doc` rc 1 naming that one variable; with the entry, rc 0 over 396 scanned names. **This is [#2312](https://github.com/mudler/vllm.cpp/issues/2312) recurring, the same class within one day** -- that row records the identical failure for `VT_DFLASH_BOUNDS_DEVICE` from [#2304](https://github.com/mudler/vllm.cpp/pull/2304), fixed by [#2313](https://github.com/mudler/vllm.cpp/pull/2313) with the same note that it is a base failure every later branch inherits. Two occurrences in a day suggests the gap is STRUCTURAL rather than an oversight: nothing forces the doc entry at the point the knob is introduced, and the gate that would catch it only runs against a base that already merged. **Documented in `docs/ENVIRONMENT.md` rather than allowlisted**, mirroring #2313 and the convention this knob's own family sets -- `VT_QWEN35_ALIAS_HOST_WEIGHTS`, whose behaviour this variable governs, is documented there, and the allowlist is for kernel-internal tuning switches. It is user-facing on its face: it decides whether a dense weight is staged as a true device copy or left aliased, which is the difference #1299 measured between a model that decodes and one that exhausts a 119.631 GiB box. Semantics read from the code rather than transcribed: stage only while `free - bytes >= frac * total`, default `0.55`, with unset, empty, unparsable, `<= 0` and `>= 1` all falling back to `0.55` | bug | +| [#2261](https://github.com/mudler/vllm.cpp/issues/2261) | `MODEL-MM-QWEN4-EXP` | **The G4 llama.cpp ladder cannot run: `llama-server` at the `qwen4exp` pin reports NO KV size, so `KV_BYTES_PER_TOKEN` has to be measured on a lease.** `scripts/qwen4exp-llamacpp-ladder.sh` extracted `KV self size = N` from the server log and, finding nothing, set the term to 0 and passed. Measured on the row's own production capture — `decode-proof/llama-server.log`, the COMPLETE unfiltered server output at 1,862 bytes — there is no `KV self size`, no `llama_kv_cache:` sizing line and no allocation summary at all, with sixteen minutes between `load_model:` and `threadpool init` and nothing printed in between; `/props` carries no KV bytes either, its only sizing fields being `n_ctx = 4096` and `total_slots = 1`. `KV_BYTES_PER_TOKEN` also defaults to 0, so on the real box NEITHER check carried a KV term while the ladder configures `CTX_TOTAL=49152` over 32 slots against a 67.5 GiB model on a 119 GiB unified-memory device that reboots rather than swaps — the 128-GiB-on-a-119-GB-box family, wearing the guard's own name. Nothing caught it because every server stub in the test suite emitted the line, so the fixture and the measured denominator disagreed. The guard now REFUSES (`E_KV_UNREPORTED`, 21) when neither the engine nor the operator supplies a term, and the fixture is silent about KV as the real server is. `/metrics` was read and rejected (it publishes a usage RATIO, not a size); a post-launch RSS check was rejected because whether GB10's unified `cudaMalloc` shows in `smaps_rollup` cannot be settled without a lease. Owed: a per-token cost measured under `-np 32 -c 49152` on a leased load | bug | +| [#2262](https://github.com/mudler/vllm.cpp/issues/2262) | `MODEL-MM-QWEN4-EXP` | **The llama.cpp arm's mutation sweep is not re-executable and its CUDA toolchain is asserted rather than pinned.** Two reproducibility debts narrowed rather than closed by the repair that brought `scripts/qwen4exp-llamacpp-build-cuda.sh` and `scripts/qwen4exp-llamacpp-decode-proof.sh` into the tree. (1) `docs/bench-evidence/qwen4exp-llamacpp-ladder-arm-20260829.md` records 11 mutations red with none unarmed, and the sweep DRIVER is not committed, so nine of them are a claim about a run that happened once on one machine; two are executable tests, `test_set_u_mutation_removing_one_default_goes_red` and `test_kv_mutation_restoring_the_fail_open_default_goes_red`. (2) `apt-get install -y cuda-toolkit-13-0` pins a CHANNEL, not a version, so a rerun gets whatever apt serves and, before this, nothing would have noticed; the build now carries `EXPECT_NVCC=13.0.88`, the version the evidence records, compares it against `nvcc --version` and exits 89 on a mismatch, which makes drift visible without making apt serve one version. Owed: a committed sweep driver (or the nine as tests), and a genuinely pinned toolchain — a versioned apt pin or a recorded container image | record | +| [#2336](https://github.com/mudler/vllm.cpp/issues/2336) | `MODEL-MM-QWEN4-EXP` | **The layer loop's remaining prerequisite is the PLE BLOCK and its GATE op, and the spec's `## Now` contradicted itself about that on `c0fa299b1` — two LIVE enumerations, which is #2288 in its seventh turn.** One paragraph said "NONE remain. The count is ZERO" of #2249's five prerequisites (true, and only about those five); eleven lines below, "What has no production shape yet is the PLE block, the GDN weight adapter onto `GdnLayerWeights`, the hyper-connection stream through the per-layer loop, and the loop itself". A wave dispatched to write the loop read the first and returned `NEEDS_DECISION`. Measured on `bd90b92b0`, the count moves BOTH ways. The PLE GATE (`modeling_qwen4_exp.py:1181-1182` — `gate.abs().clamp_min(1e-6).sqrt() * gate.sign()`, then `sigmoid(gate) * value.unsqueeze(-2)`) was OP-SIZED and nothing had ever named it: a `git grep` for `clamp_min`, `signed_sqrt` and `copysign` over `src/vt include/vt` returned ZERO lines and the only implementation was the host, file-local `SignedSqrtGate`. The DOT around it needs no new op (`vt::BatchedMatmul` over `[T*hc,1,H] x [T*hc,H,1]` VIEWS), and the multiply cannot reuse `vt::SigmoidGateBf16` (refuses by element count) or `vt::MulColVecF32` (per-COLUMN) because BOTH its operands broadcast — so ONE fused op is owed, not five. Two listed items are smaller than "missing production shape" implies: the GDN adapter is a nine-assignment FIELD COPY (with `output_gate_type` sigmoid-vs-silu, a lost `ResidentWeight::d_dev`, and an empty `in_proj_ba` as its real risks) and the hyper-connection widen is `vt::IndexSelect`. What is left is the PLE BLOCK — the LAST block seam, `PleForward` has zero cross-TU callers and `vt::RmsNormGroup` has zero production callers — and the loop. Split W5e-1 (the gate op), W5e-2 (the block), W5f (the loop). **W5e-1 LANDED: `vt::Qwen4ExpPleGate`, CPU arm, gated against section J of `qwen4_exp_ple_goldens.inc`, UNREACHED by design and recorded under `## Owed`; W5e-2 and W5f remain open on this issue.** Spec [qwen4-exp-flash-next.md](../specs/qwen4-exp-flash-next.md) | bug | +| [#2337](https://github.com/mudler/vllm.cpp/issues/2337) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **W5b-2b — the weight bridge for the other four arms and the engine binding: `ModelRegistry::Forward` REACHES GLM-5.3-Flash.** Split out of [#2241](https://github.com/mudler/vllm.cpp/issues/2241) for the same reason W5b-1 was split into [#2324](https://github.com/mudler/vllm.cpp/issues/2324): #2241's index row is spent on W5b-2a and this file is append-only with one row per issue. `ForwardGlm5NextForConditionalGeneration` stops refusing by name, which discharges the reachability halves of O15, O16, O17, O23, O25 and O26 — the six debts this row has carried since W2 saying the KDA arm, the mHC bricks, the DSA indexer, the MoE block, W5b-1's attention plus bridge and W5b-2a's decoder layer were each gated and none reached from a production entry point. **The MoE half was a RESIDENCY problem and not plumbing**: one sparse layer's three expert banks are 27.0 GiB in f32 and the 42 sparse layers together are 1,134 GiB against ~119.63 GiB usable, 9.5x over, and `kBridgeTensorF32ByteCeiling` already refused one 9.0 GiB bank by name — which is that gate working, not an obstacle. It did NOT move. `num_experts_per_tok` is 8 of 288, so `MoeLayerWeights` grows a borrowed `ExpertSource*` and `DecodeOwnedTensorRowsToF32` decodes a contiguous leading-axis ROW RANGE out of a block-resident tensor: one expert is 100,663,296 f32 bytes (0.09375 GiB) and one bank row is 33,554,432 — **32x UNDER the same unchanged ceiling the bank is 9x over** — and the RANGE is checked against that ceiling, so asking for all 288 rows is refused by exactly the arithmetic that refuses the whole tensor. A block row that is not a whole number of blocks is refused by name, because a mid-block slice does not fail: the decoder reads the next block's scale and returns plausible values from the wrong quantization. `MoeForward` now visits each HIT expert ONCE, grouped, which is upstream's own order (`Glm5NextTextExperts.forward` loops the hit experts, not the tokens) and what bounds the peak at one expert; every `[t, j]` slot is still computed independently, so the resident path is byte-identical and is asserted EXACTLY against a bank-resident reference. The same row range serves the two other tensors no device here holds in float — `token_embd.weight` and `output.weight`, 2.36 GiB each — as a per-token gather and a 64 MiB-chunked head. **The binding** is `glm5_next_forward.{h,cpp}`: `Glm5NextGgufLayerSource` holds ONE layer slot and drops the previous layer before bridging the next, and `TextModelForward` grows an overload over that source so the manifold, the `prev_topk` threading and the `hc_head` collapse stay in one loop rather than being copied. The hook follows the surveyed house pattern — `NemotronHForCausalLM` and `KimiLinearForCausalLM` both ignore the paged caches and re-run the whole prefix, and no `LoadedModel` in this tree keeps per-request state — with ONE divergence in the safe direction: both precedents take `token_ids` as one sequence whatever `num_reqs` says, which attends across the request boundary, so a multi-request step is REFUSED BY NAME and ragged batching is owed. A non-CPU queue is refused too, because every primitive on this row is host f32 and `vt::MoeRouterTopK` dispatches on the queue's device. **RED FIRST from an EXISTING gate**: `test_glm5_next_scaffold`'s "the forward REFUSES BY NAME" case went red at 8 assertions the moment the hook ran, and the pin MOVED with the change the way W3's `MlaBlockDims` pin moved rather than being deleted by it. **Thirteen negative mutations on one tree, ALL thirteen killing their gate — after TWO SURVIVED the first suite and were repaired in the same branch.** The reachability mutation is the deliverable: deleting the `Glm5NextHostForward` call in the registry hook reds `test_glm5_next_forward` at 11 of 118. M12, swapping the two mHC sites in the layer source, left every logit BIT-IDENTICAL because the fixture's mHC ramps saturate every sigmoid and the Sinkhorn projection converges identically from either site — a gate that could only see it through the logits is a mute switch at that geometry, so the mapping is now asserted STRUCTURALLY with the two tensors asserted to DIFFER, and it reds 24. M13, removing the per-expert grouping, survived because the case ran ONE token, where every selected expert is hit once whatever the code does; a six-token case now fills 12 slots from 2 distinct experts and it reds 3. M5 kills by SIGSEGV (rc=139) rather than by an assertion, which corrected the refusal's own message — without it the loop dereferences a null source, it does not read zeros. Green: forward 9/118, bridge 19/32228, scaffold 38/2652, six sibling suites unchanged and green. **No token, load or speed number is claimed for the 321.32B model and none was gated** — CI runs a synthetic 4-layer miniature at `hidden_size` 32, O1 stands, and O27 names what is still owed: the `shared` indexer arm, W5b-2a's `LayerCache` binding (the full-prefix recompute does not call it), ragged batching and the device arm. Campaign [#1998](https://github.com/mudler/vllm.cpp/issues/1998) | feature | +| [#2343](https://github.com/mudler/vllm.cpp/issues/2343) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **GLM-5.3-Flash LOADS on `dgx:gpu0` and the ENGINE refuses ABOVE the model's own forward, so W5b-2b's product claim of generation is FALSE and is corrected here.** Measured 2026-08-30 by driving the production C ABI at the staged 101.2535 GiB `unsloth/GLM-5.3-Flash-GGUF UD-Q2_K_XL` artifact — `vllm-cli --device cpu --max-tokens 8` on a `vllm-cli` built at `349df8e9a` in the leased container. **The load SUCCEEDED**, which is the FIRST materialized load of this model that has ever happened: all four shards open, the tower materializes, and the engine sizes its caches — `max_model_len` auto-fits from 1048576 to 8192 against 256 blocks of 32 tokens, and `max_num_seqs` drops from 32 to 1 because one 4,390,912-byte GDN state fills a unified page of 4288 tokens. **The FIRST step then threw**, at the `input.multi_kv` guard at the TOP of `ModelRegistry::Forward`: `22 KV cache(s) from 2 published group(s) reached this forward, first 'model.layers.3.self_attn.attn', with block tables gathered for 3 of 3 published group(s), and no registered forward consumes a cache set keyed by layer name`; `vllm-cli: completion failed (status 3)`. **NO TOKEN WAS GENERATED.** That guard is KV-DSV4-MULTICACHE W3's ([#2068](https://github.com/mudler/vllm.cpp/issues/2068)) and fires for ANY model publishing a multi-cache topology BEFORE dispatch to its hook; GLM-5.3-Flash publishes three groups (W5, [#2223](https://github.com/mudler/vllm.cpp/issues/2223)), so the engine stops there and the consuming forward is that row's to write, not this one's. **What it invalidates:** W5b-2b ([#2337](https://github.com/mudler/vllm.cpp/issues/2337)) landed "LOADS AND FORWARDS on `--device cpu`" in `docs/FEATURES.md` and "A `glm5next` file LOADS and FORWARDS" in `docs/USAGE.md`, both written from the focused gate — which is exactly the reading this measurement corrects — and both are repaired in the same change, because a record correction that leaves the lie in product output is not a correction. O27's reachability discharge STANDS in the letter it was made in (the entry point dispatches to the hook when `multi_kv` is null, and deleting the call site still reds the gate at 11 of 118) and must NOT be read as "a user can generate text". **What is newly true:** a materialized load exists, correcting the first clause of O7 and of the `docs/USAGE.md` weights row; load plus engine init took under 26 minutes wall, which is a DURATION and not a throughput number. **Peak RSS is still owed and was NOT sampled** — a defect in the staging script, not a property of the run. Two further staging defects are recorded rather than hidden, because each produced a job that looked like a product result: `/usr/bin/time` is absent in the leased container (rc=127, the build never started) and a reused `/tmp/b` held a cmake cache keyed to the previous attempt's source path. Spec `## Owed` O28, campaign [#1998](https://github.com/mudler/vllm.cpp/issues/1998) | bug | +| [#2345](https://github.com/mudler/vllm.cpp/issues/2345) | `ENG-HYBRID-PLACEMENT` | **`hybrid-placement.md` claimed `RunMoeBlockPlaced` executed under `test_placed_moe_roundtrip` "byte-identical to the direct call and mutation-proven", after both had been deleted -- so the spec advertised a byte-for-byte placement gate that does not exist.** `866075b2f` ([#2309](https://github.com/mudler/vllm.cpp/issues/2309)) deleted the helper once W3c had made it dead; `6416aab85` ([#2331](https://github.com/mudler/vllm.cpp/issues/2331)) deleted the test, which was calling a symbol that no longer existed and had stopped `main` building. **The claim cannot simply be rewritten, for a structural reason:** `RunMoePlaced` short-circuits when the placement device equals the engine device (`if (placed_on == engine_device) return body(engine, dh);` -- no copy, no allocation), so the transfer path is reachable ONLY cross-device. The deleted test reached it by passing `kCPU` as the placement device explicitly, which the seam no longer accepts. FIXED by correcting the bullet to say there is no such gate and pointing at the cross-device gate recorded under `## Owed` in [expert-stream-device-slots.md](../specs/expert-stream-device-slots.md). **Filed rather than quietly edited, and this is the point of the row:** it is the fallout of #2331, which was my own change, and a record that OVERSTATES coverage is precisely the defect that cost real time hours earlier the same day -- [#2302](https://github.com/mudler/vllm.cpp/issues/2302), a wrong dependency written into a spec because two stale records agreed with each other and neither was the tree. Same shape, same treatment: a traceable correction rather than a silent one. The narrative at `hybrid-placement.md:442` is accurate HISTORY of how the code got here and is deliberately untouched; only the live-coverage claim was wrong | bug | +| [#2282](https://github.com/mudler/vllm.cpp/issues/2282) | `BACKEND-TENSTORRENT-QWEN35` | **The residency state manufactures the staging writes it then pays for: 7-8 full-tensor mesh-CQ writes per decode step, each charged the per-op CQ tax W5 measured.** The W6 probe (#2273) counted 30 persistent-route restages on a 3-token eager leg — `[11,6144]`×17 (one stable activation-hidden slot) + `[176,128]`×13 (three rotating pool bases) — and the causal chain is entirely ours: `MarkHostWritten` (`tenstorrent_ops.cpp:5654`, callers `tenstorrent_backend.cpp:56,66,70`) marks a slot host-current/device-stale, including `OnScratchBlockAcquired` where a retained DevicePool block becomes a NEW tensor whose device bytes are garbage; `CommitHost` (`:1231`) drops the device shadow entirely on a host in-place write; the next device use re-uploads the FULL tensor through the persistent arm (`:561`). The lever: eliminate staging writes instead of amortizing them — a residency state precise enough that a step stages each slot's bytes once, or zero times when the consumer overwrites the full buffer on device (candidate mechanisms: device-will-overwrite reservation, narrowed `CommitHost`, upload-on-write; the implementer derives and records the actual mechanism with explicit restage semantics per the W5 aliasing awareness). Invariant: staging stays bit-identical — sacred golden pair 16/16 STRICT, full TT suite green, speed never tokens; capture-unsafe refusals keep semantics; f32 arms keep declared dtypes; #1486 never-destroy holds; the production decode path's write count must move, observably. Evidence owed: same-method before/after on the P150 (identical leg, JIT-discard, one lock hold) reporting BOTH per-step write count and wall time, plus a fresh benchmark-record entry; a count that does not drop or a wall that does not move is a reported result — the attribution shifts or the lever is named unreachable with the trace that proves it. Owned by `BACKEND-TENSTORRENT-QWEN35`, successor to W6 (#2273, closed as inexpressible via #2280) | perf | +| [#2294](https://github.com/mudler/vllm.cpp/issues/2294) | `BACKEND-TENSTORRENT-QWEN35` | **`CopyDeviceDeviceIfCapture` records src's geometry in dst's slot without updating `dev_rows`/`dev_cols` — latent until a differing-geometry D2D copy exists.** Found by the W7 fresh review ([#2282](https://github.com/mudler/vllm.cpp/issues/2282)), audited statically at the repair head: the arm (`tenstorrent_ops.cpp:5948`) replaces dst's device shadow with a clone of src's device tensor — src's logical shape — without updating the slot's recorded geometry, and its guard only requires equal slot byte sizes (dtype-blind), so two tracked slots of equal byte size but different `[rows,cols]` geometry (or different element sizes) would make a later exact-shape `EnsureDevice2D(dst, recorded_rows, recorded_cols)` hit return src-shaped bytes to a declared-geometry consumer. No such caller exists today, so the defect is latent: every `qwen3_5.cpp` copy site (569, 1185, 1230, 1298, 1308, 1380-1389, 1467) passes a host source and is refused at `FindSlot(src) == nullptr`; the only tracked device→device callers are GatherRows-style row gathers (`qwen3.cpp:216`, `commandr.cpp:176`, `deepseek_v2.cpp:642`, the gemma/commandr family) whose destination is allocated with the source's dtype and row geometry, so slot byte equality implies equal logical geometry. Fix when it goes live: the same one-liner W7 applied to `CopyDeviceDeviceIfResident` — set `dev_rows`/`dev_cols` from `src_dev.logical_shape()` in the second lock scope — with a bit-identical audit or a test proving the differing-geometry case, since the arm runs under capture where staging semantics are strictest. Listed in the row spec's `## Owed` | bug | +| [#2353](https://github.com/mudler/vllm.cpp/issues/2353) | `KV-DSV4-MULTICACHE` | **`ModelRegistry::Forward`'s multi-cache refusal named ONE owner where THREE architectures now arrive, and never named the arriving one.** The string ended `(row KV-DSV4-MULTICACHE W5 owns the consuming forward; #1925, #2068)`, which was true by construction when W3 wrote it because DeepSeek-V4 was the only thing that could publish a multi-cache topology. Three reach it now: `DeepseekV4ForCausalLM` (7 groups, all attention; W5 does own it), `Qwen4ExpForConditionalGeneration` (3 groups, owned by `MODEL-MM-QWEN4-EXP`; `Qwen4ExpTextModel::Forward` does not exist) and `Glm5NextForConditionalGeneration` (3 groups, owned by that model's own row). That row's W5 is scoped in [kv-dsv4-multicache.md](../specs/kv-dsv4-multicache.md) `## Work breakdown` as the DeepSeek-V4 DSA-sparse path that removes `deepseek_v4.cpp`'s `(void)attn_kv`, so the clause was FALSE for two of the three. **The cost is measured**: [#2343](https://github.com/mudler/vllm.cpp/issues/2343) drove GLM-5.3-Flash on `dgx:gpu0` on 2026-08-30 and stopped at this guard, and its index row, `docs/FEATURES.md`, `docs/USAGE.md` and [CLAIM-GLM53-FLASH-W5B2B.md](../claims/CLAIM-GLM53-FLASH-W5B2B.md) then each had to reconstruct in prose what the string should have said. **The repair names the architecture and computes it** from `model.registration().architecture`, the handle this function already holds, and does NOT enumerate the three rows: a hard-coded list in a refusal is exactly the construct [#2288](https://github.com/mudler/vllm.cpp/issues/2288) has driven stale six times on the sibling row, in both polarities. Two stale anchors are repaired with it -- the comment cited `deepseek_v4.cpp:2886-2887,:2959-2960` for the discarded `attn_kv` and the values are `:3033-3034,:3105-3106`, 147 lines on -- and the ownership pin in `tests/vllm/v1/worker/test_runner.cpp` MOVED from `KV-DSV4-MULTICACHE W5` to `W3` rather than being deleted, with a negative assertion that the old clause is gone. **The guard is NOT lifted and the reason is measured**: all three arriving forwards would discard the caches (`deepseek_v4.cpp:3033-3034,:3105-3106` and `glm5_next_registry.cpp:156-157` are literal `(void)`, `qwen4_exp_registry.cpp:142` refuses unconditionally), so lifting it trades a refusal for a silent full-prefix recompute. Owned by `KV-DSV4-MULTICACHE`, FIXED IN FLOW and closed by this change; the lift and the by-name channel's recurrent gap are carried under that spec's `## Owed` | bug | + diff --git a/.agents/issues/GATE-ISSUE-ARCHIVE-RESTORE/ISSUE-LOCAL-01M3NR65JDEV0NH28W5ZPQDPYC.md b/.agents/issues/GATE-ISSUE-ARCHIVE-RESTORE/ISSUE-LOCAL-01M3NR65JDEV0NH28W5ZPQDPYC.md new file mode 100644 index 000000000..ff462fb4d --- /dev/null +++ b/.agents/issues/GATE-ISSUE-ARCHIVE-RESTORE/ISSUE-LOCAL-01M3NR65JDEV0NH28W5ZPQDPYC.md @@ -0,0 +1,19 @@ +ID: ISSUE-LOCAL-01M3NR65JDEV0NH28W5ZPQDPYC +Title: the frozen archive dropped 27 rows its own evidence records still cite +Row: GATE-ISSUE-ARCHIVE-RESTORE +State: OPEN +Kind: bug +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-29 +Updated: 2026-09-29 +Closed: - + +## Problem + +Commit 2e84a073b deleted .agents/completed/issue-index.md wholesale (913 lines) and e3539d994 restored a hand-picked 886-line copy that dropped 27 archived rows. Those 27 rows are exactly what 43 frozen-evidence records cite: 24 cite a line past the restored file's end, 19 cite a line that now holds a different row. No gate sees any of that today: the frozen-evidence comparison runs only in the _intake branch, whose 9 records all cite surviving lines, so all 43 stale quotes are row-owned records that validate untouched. But every one of those 43 quotes is byte-true to the pre-deletion archive, so no edit to any record can repair them against a truncated archive; the defect is in the archive, not in the quoting. Measured: re-inserting the deleted lines at the positions difflib reports leaves the file byte-identical to the pre-deletion archive except line 406, which keeps e3539d994's own deliberate link re-point; the 831-quote census moves from 350 byte-equal to 373, and check-agent-record stays green with unchanged row counts. With the #3350 relative-link comparison the restored archive resolves all 831 quotes, 0 failing. Fix: splice the 27 rows back at their original positions, re-pointing any link that does not resolve from .agents/completed/ (measured: none needed). Prerequisite for the stacked checker change that enforces the comparison outside _intake. + +## Resolution + +- diff --git a/.agents/issues/GATE-ISSUE-ARCHIVE-RESTORE/ISSUE-LOCAL-01M3NSTDJSHREP8HCX83K5KXV0.md b/.agents/issues/GATE-ISSUE-ARCHIVE-RESTORE/ISSUE-LOCAL-01M3NSTDJSHREP8HCX83K5KXV0.md new file mode 100644 index 000000000..e2f496caa --- /dev/null +++ b/.agents/issues/GATE-ISSUE-ARCHIVE-RESTORE/ISSUE-LOCAL-01M3NSTDJSHREP8HCX83K5KXV0.md @@ -0,0 +1,19 @@ +ID: ISSUE-LOCAL-01M3NSTDJSHREP8HCX83K5KXV0 +Title: the frozen-evidence contract is not enforced outside _intake +Row: GATE-ISSUE-ARCHIVE-RESTORE +State: OPEN +Kind: bug +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-29 +Updated: 2026-09-29 +Closed: - + +## Problem + +validate_issue_record runs the frozen-evidence comparison (the quoted archive line must be byte-equal to the declared line, and must name the record's own GitHub number) only in the _intake branch. Row-owned and _owed records that carry a Frozen archive evidence block are never compared: any of the 822 non-intake blocks could drift from the archive, swap a URL, or quote another issue's row, and every gate stays green. Measured on the restored archive (#3351 data half): 831 records carry a block, 831 resolve under the #3350 comparison with the record directory as base, 831 quote a line carrying their own issue number, 0 violations, so enforcing the same rule outside _intake adds no new red today while closing the drift door. Absence of the block stays legal outside _intake (456 row-owned records legitimately have none); presence is not. Fix: hoist the comparison into a shared helper and apply it in the _owed and row-owned branches, and strip a trailing CR from the archived line so a CRLF Windows working copy compares equal to the LF committed blob. + +## Resolution + +- diff --git a/.agents/issues/POLICY-ISSUE-INTAKE/ISSUE-LOCAL-01M3NH23E78PEXCJF4HW1XHQQ3.md b/.agents/issues/POLICY-ISSUE-INTAKE/ISSUE-LOCAL-01M3NH23E78PEXCJF4HW1XHQQ3.md new file mode 100644 index 000000000..156429915 --- /dev/null +++ b/.agents/issues/POLICY-ISSUE-INTAKE/ISSUE-LOCAL-01M3NH23E78PEXCJF4HW1XHQQ3.md @@ -0,0 +1,19 @@ +ID: ISSUE-LOCAL-01M3NH23E78PEXCJF4HW1XHQQ3 +Title: frozen-evidence byte equality and record link resolution cannot both hold +Row: POLICY-ISSUE-INTAKE +State: OPEN +Kind: bug +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-28 +Updated: 2026-09-28 +Closed: - + +## Problem + +An _intake record's Problem must quote the frozen archive line byte for byte, but check-agent-record's check_links requires every link in a record to resolve from the record's own directory. The archive lives at .agents/completed/issue-index.md, one level under .agents, so a link to a spec is spelled ../specs/x.md there; the record that QUOTES the row lives at .agents/issues//, two levels down, so the same link must be spelled ../../specs/x.md from there. One string cannot satisfy both. Commit e3539d994 re-pointed the ISSUE-GH-1033 quote to the record-relative spelling to fix the dangling link, which broke the byte comparison: first divergence at column 2191 of archive line 350, archive 2237 chars against evidence 2240. Measured over all 831 records that carry a Frozen archive evidence block, 350 are byte-equal, 438 differ from the cited line ONLY by a relative link rebase, and 43 cite a line the archive no longer has (24 past its 886 lines after the delete and re-add, 19 naming a row that moved). Restoring the archive spelling makes check-links red instead. Fix: compare the quote against the cited line with each side's relative link targets resolved to the file they denote, so the check asks whether the quote IS the archived row rather than which directory holds the quote. + +## Resolution + +- diff --git a/.agents/specs/gate-issue-archive-restore.md b/.agents/specs/gate-issue-archive-restore.md new file mode 100644 index 000000000..b605db193 --- /dev/null +++ b/.agents/specs/gate-issue-archive-restore.md @@ -0,0 +1,83 @@ +# Spec — the frozen archive dropped 27 rows its own evidence records still cite + +Row: `GATE-ISSUE-ARCHIVE-RESTORE` (unplaced record/gate defect; the completed +archive is a record surface, not a matrix row) +State: `ACTIVE` + +## Scope + +`.agents/completed/issue-index.md` is the frozen archive every +`### Frozen archive evidence` block quotes. It is not frozen in the way the +name claims: + +1. Commit `2e84a073b` deleted the file wholesale — 913 lines — alongside its + (legitimate) spec addition. +2. Commit `e3539d994` restored a hand-picked copy: "The last pre-retirement + revision is restored at `.agents/completed/issue-index.md`, with its + internal links re-pointed." The restore measured 886 content lines: 27 + archived rows came back missing, and one surviving line (406) was + deliberately re-pointed. + +The 27 dropped rows are not idle history. Exactly 43 records under +`.agents/issues/` cite them: 24 cite a line past the restored file's end, 19 +cite a line that now holds a different row. No gate sees any of that today: +the frozen-evidence comparison runs only in the `_intake` branch of +`validate_issue_record`, whose 9 records all cite surviving lines, so all 43 +stale quotes are row-owned records that validate untouched. But every one of +those 43 quotes is byte-true to the pre-deletion archive, so no edit to any +record can repair them against a truncated archive. The defect is in the +archive, not in the quoting. + +Measured over all 831 records that carry a frozen-evidence block, against +the committed LF blob: + +- before the restore: 350 byte-equal, 458 failing under this base's + byte-only comparison (481 under the #3350 link-rebase comparison); +- after re-inserting the 27 lines at the positions `difflib` reports + between `2e84a073b~1` and the restored file: the file is byte-identical + to the pre-deletion archive except line 406, which keeps `e3539d994`'s own + deliberate re-point; 373 byte-equal, and under #3350's comparison all 831 + resolve, 0 failing. + +`scripts/check-agent-record.py` stays green with unchanged row counts +(ENGINE=179 MODEL=384 QUANT=87 KERNEL=60 BACKEND=90); none of the re-pointed +links in the restored lines dangle from `.agents/completed/`, so nothing +needed re-pointing beyond what `e3539d994` already did. + +In scope: + +1. Splice the 27 dropped rows back at their original positions. +2. Re-point any restored link that does not resolve from + `.agents/completed/` (measured: none). +3. The red-before / green-after census in the commit message. + +Out of scope, each for its own reason: + +1. **Editing any record.** The 43 quotes are correct; "repairing" them would + re-anchor records onto a truncated archive and make the falsification + load-bearing. +2. **Changing any checker.** Widening or narrowing a gate to make a red go + green is what AGENTS.md forbids; the archive, not the gate, is wrong. +3. **Enforcing the frozen-evidence comparison outside `_intake`.** That is + the checker half of this pair and lands as its own stacked PR (#3350 is + the comparison it needs; this restore is the data it needs). +4. **The 4 row-cell anomalies in the ROW bucket** — 3 records whose archived + cell is the em-dash placeholder (`ISSUE-GH-83`, `ISSUE-GH-606`, + `ISSUE-GH-408`) and 1 genuine cross-row citation (`ISSUE-GH-298` cites a + `PERF-27B-LMHEAD-DSR` line while living under `PERF-27B-LMHEAD-FP4`). + They are pre-existing record-content questions, orthogonal to line + existence, and every one of them still resolves after the restore. + +## Enforcement (the stacked checker PR) + +With the rows restored and #3350's relative-link comparison landed, the +frozen-evidence contract is enforced wherever the block appears, in every +owner directory: a record that QUOTES an archived row must quote the line it +declares, modulo exactly the relative-link rebase the record's directory +forces, and the quoted line must carry the record's own GitHub number. +Absence of the block stays legal everywhere except `_intake`; presence is +not. Measured over the corpus on the restored archive: 831 records carry a +block, 831 resolve, 831 identify their own issue, 0 violations -- the +ratchet adds no new red. Working-copy EOL no longer changes the answer: the +comparison strips a trailing CR from the archived line, because the +committed blob is LF and a Windows checkout is not. diff --git a/scripts/issue_records.py b/scripts/issue_records.py index 544f34632..d3390b186 100755 --- a/scripts/issue_records.py +++ b/scripts/issue_records.py @@ -13,6 +13,7 @@ from dataclasses import dataclass from datetime import date from pathlib import Path +import posixpath import re from typing import TypeAlias @@ -38,11 +39,21 @@ _LOCAL_ID = re.compile(r"ISSUE-LOCAL-([0-7][0-9A-HJKMNP-TV-Z]{25})\Z") _ROW_ID = re.compile(r"[A-Z0-9][A-Za-z0-9_.-]*\Z") _GITHUB_NUMBER = re.compile(r"[1-9][0-9]*\Z") +# The one frozen archive an intake record's Problem must quote, in +# repository-relative form. The record that HOLDS a quote lives deeper than +# the file the quote was cut from, which is why the comparison below +# resolves relative link targets instead of comparing their spelling: the +# same row is `../specs/x.md` in the archive and `../../specs/x.md` in the +# record, and both gates (this one and check-links) must be satisfiable at +# once. +FROZEN_ARCHIVE_RELPATH = ".agents/completed/issue-index.md" _INTAKE_PROBLEM = re.compile( - r"Archive: `\.agents/completed/issue-index\.md:([1-9][0-9]*)`\n\n" + r"Archive: `" + re.escape(FROZEN_ARCHIVE_RELPATH) + r":([1-9][0-9]*)`\n\n" r"### Frozen archive evidence\n\n" r"> ([^\n]+)\Z" ) +_LINK_TARGET = re.compile(r"(\[[^\]]*\]\()([^)]+)(\))") +_REMOTE_TARGET = re.compile(r"[A-Za-z][A-Za-z0-9+.-]*:") class IssueRecordError(ValueError): @@ -303,11 +314,67 @@ def intake_archive_evidence(record: IssueRecord) -> tuple[int, str] | None: return int(match.group(1)), match.group(2) +def _resolve_relative_links(line: str, base: str) -> str: + """Rewrite every relative Markdown link target to the file it denotes. + + A relative link resolves against the file that QUOTES it, so one frozen + row cannot keep one spelling in the archive (`.agents/completed/`) and in + the record that quotes it (`.agents/issues//`). Resolving both + sides to the repository-relative path they point at is what lets the + evidence comparison ask the question it means to ask -- is this the + archived row? -- instead of which directory holds the quote. + + Remote (`https:`, `mailto:`, ...) and root-absolute (`/...`) targets are + left exactly as written: only a rebase of a relative target is + comparable, so a swapped remote URL, a truncated link, or any other + difference still fails the comparison byte for byte. + """ + + def rewrite(match: re.Match[str]) -> str: + raw = match.group(2) + lead = raw[: len(raw) - len(raw.lstrip())] + trail = raw[len(raw.rstrip()) :] + target = raw.strip() + angled = target.startswith("<") and target.endswith(">") + if angled: + target = target[1:-1] + path, separator, fragment = target.partition("#") + if ( + not path + or path.startswith(("/", "?")) + or _REMOTE_TARGET.match(path) is not None + ): + return match.group(0) + resolved = posixpath.normpath(posixpath.join(base, path)) + rendered = f"{resolved}{separator}{fragment}" + if angled: + rendered = f"<{rendered}>" + return f"{match.group(1)}{lead}{rendered}{trail}{match.group(3)}" + + return _LINK_TARGET.sub(rewrite, line) + + def _archive_evidence_matches_source( evidence: tuple[int, str], frozen_archive: bytes | None, + record_base: str = "", ) -> bool: - """Require exact UTF-8 evidence bytes at the declared one-based source line.""" + """Require the declared line to be the quote, modulo relative link rebase. + + Byte equality is tried first and answers almost every record: an archive + that has not moved and a record that copied the line verbatim need + nothing resolved. Only when the bytes differ is the quote compared with + each side's relative link targets resolved, which admits exactly the + spelling a MOVE forces (`../specs/x.md` vs `../../specs/x.md`) and + nothing else. `record_base` is the record's repository-relative + directory; with no base to resolve against, only byte equality passes. + + A trailing CR on the archived line is stripped before comparing: the + committed blob is LF, but a Windows working copy checks the archive out + CRLF, and byte equality must answer the CONTENT of the line, not which + checkout read it. The quote side is parsed from record text, so it has + no CR to strip. + """ if frozen_archive is None: return False @@ -315,7 +382,18 @@ def _archive_evidence_matches_source( source_lines = frozen_archive.split(b"\n") if line_number > len(source_lines): return False - return source_lines[line_number - 1] == archived_line.encode("utf-8") + source_line = source_lines[line_number - 1].removesuffix(b"\r") + if source_line == archived_line.encode("utf-8"): + return True + if not record_base: + return False + try: + decoded = source_line.decode("utf-8") + except UnicodeDecodeError: + return False + return _resolve_relative_links( + decoded, posixpath.dirname(FROZEN_ARCHIVE_RELPATH) + ) == _resolve_relative_links(archived_line, record_base) def _archive_row_owner(line: str, github: int | None) -> str | None: @@ -342,19 +420,57 @@ def _archive_row_owner(line: str, github: int | None) -> str | None: def valid_intake_archive_evidence( record: IssueRecord, frozen_archive: bytes | None = None, + record_base: str = "", ) -> tuple[int, str] | None: - """Return evidence only when exact source bytes identify ownerless self.""" + """Return evidence only when the source line identifies ownerless self. + + `record_base` is the record's repository-relative directory, so a quote + whose links were re-pointed to resolve from there still matches the line + it was cut from; see `_resolve_relative_links`. + """ evidence = intake_archive_evidence(record) if evidence is None or not _archive_evidence_matches_source( evidence, frozen_archive, + record_base, ): return None _, line = evidence return evidence if _archive_row_owner(line, record.github) in {"", "-", "—"} else None +def _quoted_evidence_errors( + evidence: tuple[int, str] | None, + frozen_archive: bytes | None, + record_base: str, + github: int | None, +) -> list[str]: + """Contract errors for a record that QUOTES an archived row. + + Wherever a record carries a Frozen archive evidence block -- _intake, + _owed, or row-owned -- the quote must be the line it declares, modulo + the relative-link rebase the record's directory forces, and the line + must be about this record's own GitHub number. Absence of the block is + legal everywhere except _intake; presence is not, in any owner + directory. Measured over the corpus at d15b1cc09 with the 27 dropped + archive rows restored (GATE-ISSUE-ARCHIVE-RESTORE): all 831 existing + blocks satisfy both halves, so the ratchet adds no new red. + """ + + if evidence is None: + return [] + if not _archive_evidence_matches_source(evidence, frozen_archive, record_base): + return [ + "Frozen archive evidence must equal the declared line in the frozen " + "archive source (a relative link may differ only by spelling, and " + "must resolve to the same file)" + ] + if _archive_row_owner(evidence[1], github) is None: + return ["Frozen archive evidence must identify this issue"] + return [] + + def validate_issue_record( record: IssueRecord, path: str | Path, @@ -455,7 +571,7 @@ def validate_issue_record( errors.append(f"{name} must be UNKNOWN for Availability METADATA_ONLY") if normalize_body(record.resolution).strip() != "-": errors.append("Resolution must be - for Availability METADATA_ONLY") - archive = ".agents/completed/issue-index.md" + archive = FROZEN_ARCHIVE_RELPATH exact_number = ( record.github is not None and re.search(rf"(? str: + """An intake Problem whose quote carries a relative link, archive-spelled. + + The archive lives at `.agents/completed/issue-index.md`, so a link to a + spec is `../specs/...` there. The record that QUOTES the row lives two + levels down, and the same link is `../../specs/...` from there -- the + spelling #3348 is about. + """ + archived = ( + f"| [#{number}](https://github.com/mudler/vllm.cpp/issues/{number}) " + f"| — | Spec [example]({spec}) | bug |" + ) + return ( + f"Archive: `.agents/completed/issue-index.md:{line}`\n\n" + "### Frozen archive evidence\n\n" + f"> {archived}" + ) + + +def linked_frozen_archive_source(number: int = 77, spec: str = "../specs/example.md") -> bytes: + archived = ( + f"| [#{number}](https://github.com/mudler/vllm.cpp/issues/{number}) " + f"| — | Spec [example]({spec}) | bug |" + ) + return ( + "# Issue index\n\nFrozen archive\n" + "| Issue | Row | Title | Kind |\n" + "|---:|---|---|---|\n" + f"{archived}\n" + ).encode() + + def parse(text: str | None = None) -> records.IssueRecord: return records.parse_issue_text(text if text is not None else issue_text()) @@ -543,6 +579,83 @@ def test_intake_evidence_must_equal_the_declared_frozen_source_line( frozen_archive=frozen_archive, ) + def test_intake_quote_may_rebase_a_relative_link_to_resolve_from_the_record( + self, tmp_path: Path + ) -> None: + """The record-relative spelling of a link is the SAME link (#3348). + + `check-links` requires every link in a record to resolve from the + record's own directory; the frozen-evidence comparison required the + quote to be byte-identical to the archive line it was cut from. One + string could not satisfy both, because the archive is one level under + `.agents/` and the record two. A quote whose ONLY difference is a + relative link resolving to the same file is accepted. + """ + record = self.intake(problem=linked_intake_problem(spec="../../specs/example.md")) + validate( + tmp_path, + record, + owner="_intake", + rows={"ROW-A"}, + owed=(), + frozen_archive=linked_frozen_archive_source(), + ) + + @pytest.mark.parametrize( + ("spec", "archive_spec"), + [ + # a rebase to a DIFFERENT file is still drift + ("../../specs/other.md", "../specs/example.md"), + ("../../specs/example.md#frag", "../specs/example.md"), + # a remote target is never rebased, so a swap must fail byte for byte + ("https://example.com/example.md", "../specs/example.md"), + # a root-absolute target is never rebased either + ("/specs/example.md", "../specs/example.md"), + ], + ) + def test_intake_quote_may_rebase_nothing_but_a_relative_link( + self, tmp_path: Path, spec: str, archive_spec: str + ) -> None: + record = self.intake(problem=linked_intake_problem(spec=spec)) + with pytest.raises(records.IssueRecordError, match="frozen archive source"): + validate( + tmp_path, + record, + owner="_intake", + rows={"ROW-A"}, + owed=(), + frozen_archive=linked_frozen_archive_source(spec=archive_spec), + ) + + def test_a_rebased_quote_needs_a_record_base_to_be_compared(self) -> None: + """Without the record's directory, only byte equality may pass. + + The public helper keeps the old strict behaviour for a caller that + cannot say where the quote lives, rather than guessing a base and + admitting a drift it cannot see. + """ + record = records.IssueRecord( + id="ISSUE-GH-77", + title="Archived 77", + row=None, + state="UNKNOWN", + kind="bug", + github=77, + mirror="MISSING", + availability="METADATA_ONLY", + created="UNKNOWN", + updated="UNKNOWN", + closed="UNKNOWN", + problem=linked_intake_problem(spec="../../specs/example.md"), + resolution="-", + ) + assert records.valid_intake_archive_evidence( + record, linked_frozen_archive_source() + ) is None + assert records.valid_intake_archive_evidence( + record, linked_frozen_archive_source(), ".agents/issues/_intake" + ) is not None + def test_full_and_local_records_cannot_use_intake(self, tmp_path: Path) -> None: for record in ( parse(), @@ -663,3 +776,159 @@ def test_body_normalization_preserves_internal_blank_lines_only(self) -> None: def test_remote_body_normalization_has_lf_and_one_final_newline(self) -> None: body = "\r\nRow: ROW-A \r\n\r\nProblem\t\r\n\r\n" assert records.normalize_github_body(body) == "Row: ROW-A\n\nProblem\n" + + +class TestFrozenEvidenceEnforcement: + """The frozen-evidence contract applies wherever the block appears. + + A record outside _intake that QUOTES an archived row must quote the row + it declares, byte for byte, modulo exactly the relative-link rebase the + record's directory forces (#3348). Absence of the block stays legal for + every owner except _intake; presence is not. + """ + + def row_owned(self, **changes: object) -> records.IssueRecord: + values: dict[str, object] = { + "id": f"ISSUE-GH-{''}91", + "title": "Archived 91, row-owned", + "row": "ROW-A", + "state": "OPEN", + "kind": "bug", + "github": 91, + "mirror": "SYNCED", + "availability": "FULL", + "created": "2026-08-01", + "updated": "2026-08-31", + "closed": "-", + "problem": linked_intake_problem(number=91), + "resolution": "-", + } + values.update(changes) + return records.IssueRecord(**values) + + def test_row_owned_quote_rebased_to_the_same_file_passes( + self, tmp_path: Path + ) -> None: + validate( + tmp_path, + self.row_owned(), + owner="ROW-A", + rows={"ROW-A"}, + owed=(), + frozen_archive=linked_frozen_archive_source(number=91), + ) + + def test_row_owned_quote_rebased_to_a_different_file_fails( + self, tmp_path: Path + ) -> None: + with pytest.raises(records.IssueRecordError, match="must equal the declared line"): + validate( + tmp_path, + self.row_owned( + problem=linked_intake_problem(number=91, spec="../specs/other.md") + ), + owner="ROW-A", + rows={"ROW-A"}, + owed=(), + frozen_archive=linked_frozen_archive_source(number=91), + ) + + def test_row_owned_quote_rebase_may_not_change_a_fragment( + self, tmp_path: Path + ) -> None: + archived = linked_frozen_archive_source( + number=91, spec="../specs/example.md#section" + ) + with pytest.raises(records.IssueRecordError, match="must equal the declared line"): + validate( + tmp_path, + self.row_owned(), + owner="ROW-A", + rows={"ROW-A"}, + owed=(), + frozen_archive=archived, + ) + + def test_row_owned_remote_target_must_stay_byte_equal(self, tmp_path: Path) -> None: + line = ( + f"| [#91](https://github.com/mudler/vllm.cpp/issues/91) " + f"| — | Spec [docs](https://example.com/guide) | bug |" + ) + archive = ( + "# Issue index\n\n" + "| Issue | Row | Title | Kind |\n" + "|---:|---|---|---|\n" + f"{line}\n" + ).encode() + problem = ( + "Archive: `.agents/completed/issue-index.md:5`\n\n" + "### Frozen archive evidence\n\n" + f"> {line.replace('example.com/guide', 'example.com/other')}" + ) + with pytest.raises(records.IssueRecordError, match="must equal the declared line"): + validate( + tmp_path, + self.row_owned(problem=problem), + owner="ROW-A", + rows={"ROW-A"}, + owed=(), + frozen_archive=archive, + ) + + def test_row_owned_quote_citing_another_issue_fails(self, tmp_path: Path) -> None: + with pytest.raises( + records.IssueRecordError, match="identify this issue" + ): + validate( + tmp_path, + self.row_owned(problem=linked_intake_problem(number=77)), + owner="ROW-A", + rows={"ROW-A"}, + owed=(), + frozen_archive=linked_frozen_archive_source(number=77), + ) + + def test_owed_quote_rebased_to_the_same_file_passes(self, tmp_path: Path) -> None: + record = self.row_owned(row=None) + validate( + tmp_path, + record, + owner="_owed", + rows={"ROW-A"}, + owed={record.id: 1}, + frozen_archive=linked_frozen_archive_source(number=91), + ) + + def test_row_without_the_block_stays_legal(self, tmp_path: Path) -> None: + validate( + tmp_path, + self.row_owned(problem="No archive quote here."), + owner="ROW-A", + rows={"ROW-A"}, + owed=(), + frozen_archive=None, + ) + + def test_no_record_base_still_requires_byte_equality(self) -> None: + evidence = linked_intake_problem(number=91) + archive = linked_frozen_archive_source(number=91) + # The record-spelled quote (what a quote CUT INTO the record looks + # like after its links resolve from the record's directory) is the + # case that needs a base: byte equality alone rejects it, because the + # archive spells the link for ITS directory. + record_spelled = evidence.rsplit("> ", 1)[1].replace("../specs/", "../../specs/") + assert records._archive_evidence_matches_source((6, record_spelled), archive) is False + assert ( + records._archive_evidence_matches_source( + (6, record_spelled), archive, ".agents/issues/ROW-A" + ) + is True + ) + # The archive-spelled quote passes byte-equality with no base at all. + archive_spelled = evidence.rsplit("> ", 1)[1] + assert records._archive_evidence_matches_source((6, archive_spelled), archive) is True + + def test_crlf_archive_source_matches_an_lf_quote(self) -> None: + line = "| [#91](https://github.com/mudler/vllm.cpp/issues/91) | — | bug |" + archive = ("# Issue index\n" + line + "\n").replace("\n", "\r\n").encode() + assert records._archive_evidence_matches_source((2, line), archive) is True