Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 3 additions & 4 deletions .agents/NOW.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,12 +22,11 @@ no per-row change needs to touch this file at all.

Token-exact (or ratified distributional) vs pinned vLLM; ≥ throughput and ≤
latency/memory on every axis, both gate models, reproduced 2–3x idle. See
[verification](verification.md). Pin: vLLM <!--pin:commit-->`e126687a9a`<!--/pin--> (<!--pin:label-->0.28.1rc1.dev132<!--/pin-->) since
2026-09-03 (#2817). **A gate HAS now run at it and it PASSED** (2026-09-04, job
[verification](verification.md). Pin: vLLM <!--pin:commit-->`a7c23ac96d`<!--/pin--> (<!--pin:label-->0.3.0.dev267<!--/pin-->), advanced
2026-09-26 from `e126687a9a` (pinned 2026-09-03, #2817) by sync `4f11dfc10`. The gate named below ran at that PRIOR pin (2026-09-04, job
`7386f034-246a-4af5-9a04-f98aafffce54`, `dgx:gpu0`, 2h15m): the OPT candidate
captured at the target is byte-identical to the committed bar --
`IDS mismatched_positions 0 of 96`, `IDS_BYTE_EQUAL True`,
`SELECTOR K=5 multi_valued_cells 0`, `TOKENGATE_VERDICT PASS`. The default
`IDS mismatched_positions 0 of 96`, `IDS_BYTE_EQUAL True`, `SELECTOR K=5 multi_valued_cells 0`, `TOKENGATE_VERDICT PASS`. The default
FLASH_ATTN backend produced the tokens, so the FA-on-GB10 risk did not fire. Our
arm's 96/96 carries over unchanged because the candidate's bytes are identical to
the bar it already passed. The BENCHMARK baselines are still measured at
Expand Down
19 changes: 19 additions & 0 deletions .agents/issues/_owed/ISSUE-LOCAL-01M3WMKTBEDAHF226YSATSR7BF.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
ID: ISSUE-LOCAL-01M3WMKTBEDAHF226YSATSR7BF
Title: five preflight doc gates red on main
Row: -
State: OPEN
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-01
Updated: 2026-10-01
Closed: -

## Problem

On origin/main the preflight carries five red record/doc gates: check-release-binary-contract and check-windows-release-state (roadmap_v1.md lost its release-lifecycle anchors and the ENG-RELEASE-WINDOWS state anchor in the record restructure), check-benchmark-index (16 detail files in docs/benchmarks/ are orphans of the BENCHMARKS.md index), check-env-doc (VT_VK_DISABLE, VT_VK_DISABLE_PAGED_ATTN, VT_VK_FENCE_TIMEOUT_MS read from src/ are neither documented nor allowlisted), and check-oracle-pins (the 2026-09-26 sync 4f11dfc10 advanced the parity pin to a7c23ac96d/0.3.0.dev267 but left the restated spans in NOW.md, FEATURES.md, how-we-measure.md, speculative-decoding.md and vllm-online-serving.md at the prior e126687a9a/0.28.1rc1.dev132). Repair the records to the checker contracts without weakening any checker.

## Resolution

Repaired on row/preflight-gate-repairs: (1)+(2) roadmap_v1.md release-lifecycle and ENG-RELEASE-WINDOWS anchors restored from 10a64f827/a0a8fedbe; (3) all 16 docs/benchmarks detail files indexed in BENCHMARKS.md; (4) VT_VK_DISABLE and VT_VK_DISABLE_PAGED_ATTN allowlisted as kernel-internal bisect hooks, VT_VK_FENCE_TIMEOUT_MS documented in docs/ENVIRONMENT.md; (5) parity pin restated to a7c23ac96d/0.3.0.dev267 in .agents/oracles/vllm.md and the five prose surfaces the 2026-09-26 sync 4f11dfc10 left behind. All five checkers exit 0; check-agent-record and the record suites stay green.
8 changes: 4 additions & 4 deletions .agents/oracles/vllm.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,11 +16,11 @@ id = vllm
role = primary
upstream = https://github.com/vllm-project/vllm
scope = every behavior vLLM implements — defaults, modes, errors, edge cases, and both correctness and speed gates
pin = e126687a9a828d513c01a07cd69f025f27d63280
pin_label = 0.28.1rc1.dev132
pinned_on = 2026-09-03
pin = a7c23ac96d7806e7c7e7d862eadbce5a33529b94
pin_label = 0.3.0.dev267
pinned_on = 2026-09-26
gateable = yes
evidence = .agents/sync/2026-09-03-e126687-runhalf.md
evidence = .agents/sync/2026-09-22-a7c23ac96d.md
```

## What this pin establishes, and what it does NOT
Expand Down
4 changes: 3 additions & 1 deletion .agents/roadmap_v1.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,7 @@
# Roadmap v1 — post-MVP

<!-- ENG-RELEASE-WINDOWS: state=ACTIVE publication=pending artifact=unpublished -->

*(user-directed 2026-07-10: this document is the live roadmap; the completed
M0-M3 record is archived at [completed/roadmap_mvp_v0.md](completed/roadmap_mvp_v0.md).)*

Expand Down Expand Up @@ -92,7 +94,7 @@ models we already ship + benchmark. Full seam map + M0–M5 W-plan:
| 14 | `ROAD-V1-D5` | LoRA, local KV/weight offload, expert streaming, wider model zoo | [engine matrix](engine-matrix.md), [model matrix](model-matrix.md) | corrected expert-streaming spike accepted (`ENG-EXPERT-STREAM` READY): bank-only safetensors→Marlin bank, fixed contiguous cache slots matching Marlin dense strides, logical→slot remap after explicit router D2H, chunked C<E prefill, whole-system GB10 memory gate; mirror floor `ENG-WEIGHT-OFFLOAD` inventoried; LoRA/local KV-offload/model-zoo unspiked | `PARTIAL` | claim expert-streaming W0 trace/c1 baseline, then W1 cache policy and W2 bank/reader/pread leaves; rest after T1 dependencies close |
| H3 | `ROAD-V1-H3` | **DIFFUSION generation — a new capability class.** MiniMax-H3 (`MiniMaxH3DiTModel`): omni-modal video+audio generation via a 50-step flow-matching denoise loop, ported from vLLM-Omni. Not autoregressive: no KV cache, sampler or logits. | [`MODEL-DIFFUSION-minimax-h3-mini-max-h3-dit`](model-matrix.md) | [minimax-h3 spike](specs/minimax-h3.md) | `PARTIAL` | **W0-W2 landed 2026-08-03**: packed layout (fp64 grid bit-exact), latent packing, scheduler and the full DiT forward parity-gated vs the upstream vLLM-Omni modules at reduced dims (max abs diff 1.6e-7, 10/10 cases). **W2b device-resident forward LANDED (f32) and GPU-VERIFIED 2026-08-03** — the whole DiT graph runs with activations resident in device memory, gated vs the same upstream goldens on a Thor sm_110 GPU at video 1.49e-7 / audio 8.94e-8. Only 3 H3 kernels were needed; the port reuses the tuned shared ops. Next gate: bf16 stream + `vt::FusedChain` glue folds, then the FP4 path — which needs sm_121a, since sm_110 resolves every fp4/cutlass feature DISABLED. **HW verdict CORRECTED 2026-08-03: e2e is NOT blocked** — quantized H3 checkpoints fit (GGUF ~41 GB working set; NVFP4 likewise) and the ComfyUI-GGUF arm's 535-tensor manifest already resolves onto our contract, so e2e + a speed comparison are reachable. W7 `/v1/videos` still needs a NEW MP4/AV-encoder dependency decision. **bf16 13-SHARD RELEASE INDEXES 2026-08-07 (`row/H3-BF16-SHARDED-DIT`)**: the ORIGINAL 66.3 GB bf16 DiT (13 safetensors shards) is now resolvable through its own `model.safetensors.index.json`, with a host-f32 reference loader and `--dit <dir>` working everywhere `--dit <file>` did; gated CPU-only (72/72, 54497 post-rebase) on index/name mapping and on the REAL 535-tensor geometry read from a sparse 13-shard release. The DEVICE streamer landed 2026-08-07 (`row/H3-BF16-SHARDED-STREAM`, spec §8.14): one tensor at a time, zero host buffer for the bulk, bit-exact vs the non-streamed reference (73/73, 55203). **ENCODER + THE NUMBER 2026-08-07 (`row/H3-ENC-BF16-COND-DIFF`, spec §8.15)**: the 14-shard bf16 text encoder streams too and `--encoder-only` runs it alone; measured over 233 tokens, Q4_K_M-vs-bf16 conditioning is cos 0.99745 mean / 6.85% rel RMS excl. sink / 3.5 deg median rotation — as much as a one-word prompt edit, but DIFFUSE. Whether the RENDER changes is NOT established (75/75, 55609). This unblocks the bf16-vs-quantized quality A/B; no render or speed number is claimed. Spec §8.13. **W-FP4a LANDED (CPU) 2026-08-06 (`row/H3-FP4-SPEED`)**: the NVFP4 DiT projections now keep FP4 PACKED and route through the shared `dense_nvfp4::MatmulNvfp4W4A16D` (Marlin W4A16 — vLLM's own forced-a16 selection; SAME kernel as Laguna/dense-Qwen3 NVFP4; no new quant code); fp4-vs-bf16 WIRING gate GREEN (62/62·30039, W4A16 dispatcher runs all 11 quantized GEMMs). **GB10 leg LANDED 2026-08-06 (`row/H3-FP4-GPU-E2E`, PR #64):** Marlin W4A16 RAN on sm_121a (`dense_gemms==11` default / `marlin_gemms==11` VT_MARLIN_DENSE=0, `fallback_gemms==0`), fp4-vs-bf16 BYTE-EXACT; fp4 is a MEMORY win (~16 vs ~66 GB), ~0.79–0.83× the bf16 arm per diffusion forward (compute-bound large M; 3.47× faster at small decode-like M). Real-checkpoint fp4-resident t2va e2e RUNS end-to-end (real 18.75 GB NVFP4 DiT + VAEs + GGUF Qwen3-VL-32B encoder → valid mp4/wav; DiT s/step 5.45/20.0/209 s @512/768/REF-768×1344-209f) but frames are a non-scene patch-grid at 12/20/50 steps → OPEN render-coherence bug (device VAE decode / denoise), separate from the fp4 speed work. vLLM-Omni serves NO quantized H3 (BF16-only) -> HW/loader-forced-indirect (4×B300 209f render 86.964 s vs 1×GB10 209 s/forward). **2026-08-08 ROW 2 DEVICE-SEAM FOLLOW-UP (#135; replaces #134):** the public 0/1 selector is mapped once to generic `DeviceType`; DSR returns 34→32 with the baseline/allowlist unchanged; CPU compile/fold test pending in CI due shared-disk pressure. |
| 15 | `ROAD-V1-D6` | **llama.cpp device breadth folded into scope (user-directed 2026-08-05):** the 11 ggml backends vLLM has no platform for — cann, musa, opencl, openvino, rpc, webgpu, zdnn, zendnn, hexagon, blas, virtgpu — inventoried as `BACKEND-GGML-*`. **SPIKES FIRST:** no implementation before each row's `.agents/specs/<slug>.md` clears the spike contract, per the standing directive. vLLM stays the mirror source; llama.cpp is the breadth reference. | [backend matrix](backend-matrix.md) | ☐ per-row spike required | `INVENTORIED` | first spike accepted |
| REL | `ROAD-V1-RELEASE` | KISS downloads per OS+host ABI: one adaptive CPU binary and one fat CUDA binary covering every supported SM; per-SM CUDA artifacts are optional diagnostics; stable channels require matching runtime evidence and build-only paths stay preview | [`ENG-RELEASE-BINARIES`](engine-matrix.md) | [release binary matrix](specs/release-binary-matrix.md) | `ACTIVE` | Required W1-W11/W13 implementation is complete in draft PR #196: eight primary bundles, extracted-archive gates, manifests/SBOM/provenance, immutable verification and attestation, generated indexes, and exact-file publication. Local CPU/Vulkan and mutation gates are green. Hosted ten-SM completion, the full eight-tuple dry run, matching-hardware evidence, merge, and tagged publication remain pending; no published binary exists. W12 remains optional/non-primary |
| REL | `ROAD-V1-RELEASE` | KISS downloads per OS+host ABI: one adaptive CPU binary and one fat CUDA binary covering every supported SM; per-SM CUDA artifacts are optional diagnostics; stable channels require matching runtime evidence and build-only paths stay preview | [`ENG-RELEASE-BINARIES`](engine-matrix.md), [`ENG-RELEASE-WINDOWS`](engine-matrix.md) | [release binary matrix](specs/release-binary-matrix.md); [Windows pre-alpha extension](specs/windows-binary-release.md) | `ACTIVE` | v0.0.2 published eight primary archive/checksum/provenance triplets plus two indexes from `7020de93652ca920424a10ac5255b34810dd2f24` in run `31466516224` (26 assets). Windows W14-W16 are implemented for one PR; native hosted gates, merged-SHA ten-tuple dry run, matching-hardware evidence, `v0.0.3-pre.1` publication and 32-asset audit remain pending. W12 remains optional/non-primary |
| IMG | `ROAD-V1-CONTAINERS` | **Published container images on GHCR, built by GitHub Actions (user-directed 2026-08-08).** The same staged bundle `ROAD-V1-RELEASE` defines, shipped from one package `ghcr.io/mudler/vllm.cpp` with the lane in the tag — `:<version>-cuda` / `-vulkan` / `-cpu`, the moving `:latest-cuda` / `:latest-vulkan` / `:latest-cpu`, and a bare `:latest` aliasing the cpu lane — and `ENTRYPOINT vllm-server`. Lanes `cuda` (one fat image, every supported SM), `vulkan`, `cpu`, plus `rocm` blocked-preview; version tags immutable, every `latest-<lane>` moves. Every lane is a `linux/amd64`+`linux/arm64` multi-arch manifest on native runners, because the project's own gate hardware (GB10, Thor, Orin) is arm64. Metal and MLX are NOT-CONTAINERIZABLE and stay static-binary-only — a recorded boundary, not pending work. Depends on the `ROAD-V1-RELEASE` install/stage tree: the image IS the bundle, so the two lanes must not grow separate layouts. No image, workflow or registry package exists. | [`ENG-RELEASE-CONTAINERS`](engine-matrix.md) | [container-images.md](specs/container-images.md) — spec ACCEPTED and **W1-W5/W7 IMPLEMENTED**: one `docker/Dockerfile` whose builder stages call the existing `scripts/build-*-release.sh`, digest-pinned ubuntu runtime bases with the CUDA runtime libs copied and the driver left to the host, ffmpeg in every lane, gated container matrix + image validator + least-privilege publish workflow | `ACTIVE` | **cpu lane BUILT AND GATED e2e 2026-08-10** (783 MB linux/amd64; `/health` 200, `/version` 200, in-container healthcheck, clean SIGTERM, booted on opt-125m). The boot gate immediately found [#312](https://github.com/mudler/vllm.cpp/issues/312): `vllm-server` ignored SIGTERM as PID 1, so `docker stop` hard-killed it (137) after 30 s — now exit 0 in 0.25 s via a self-pipe handler into the existing `server.stop()`. Two silent build-context bugs fixed on the way: `.dockerignore`'s `**/build*/` also matched FILES and was excluding `scripts/build-*-release.sh`, and the builders lacked `file`/`binutils` so the inherited archive validator failed after a full compile. **REMAINING: W6 only** — nothing is published to GHCR, cuda/vulkan are gated statically but never built here, and both arm64 legs are unbuilt so SBSA-vs-Tegra (Thor `sm_110`, Orin `sm_87`) is untouched |

An area row cannot enter `READY` without a real spike under `specs/`, and cannot
Expand Down
49 changes: 49 additions & 0 deletions .agents/specs/record-gate-repairs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# Spec: preflight record/doc gate repairs (five red gates on main)

## Scope

Repair the five red record/doc gates on `origin/main` by reconciling the
records to each checker's contract. No checker is weakened, no benchmark or
document is deleted.

1. `scripts/check-release-binary-contract.py` — `.agents/roadmap_v1.md` lost
the `| REL | ROAD-V1-RELEASE |` release-lifecycle anchors (v0.0.2 published;
Windows W14-W16 pending). Restore the historical row text (last carried at
`10a64f827`, `docs(release): add native Windows prerelease lanes (#117)`),
which is the state the checker's `EXACT_MACHINE_FIELDS`
(`archive_claims = published-v0.0.2`) and `PUBLIC_PENDING_MUTATIONS` pin.
2. `scripts/check-windows-release-state.py` — the roadmap must carry exactly
one `<!-- ENG-RELEASE-WINDOWS: state=ACTIVE publication=pending
artifact=unpublished -->` anchor plus the pending `v0.0.3-pre.1`
publication/audit phrase. Both were dropped from the roadmap in the same
restructure; restore them (anchor placement per `a0a8fedbe`).
3. `scripts/check-benchmark-index.py` — index every detail file under
`docs/benchmarks/` in the `## Benchmark detail index` table of
`docs/BENCHMARKS.md`, one row per file, ID equal to the file stem.
4. `scripts/check-env-doc.py` — classify `VT_VK_DISABLE` and
`VT_VK_DISABLE_PAGED_ATTN` (kernel-internal bisect hooks,
`src/vt/vulkan/vulkan_ops.cpp:982` and `:1296`) onto
`scripts/env-doc-allowlist.txt`; document `VT_VK_FENCE_TIMEOUT_MS` (a
user-facing operational watchdog that aborts on a wedged Vulkan dispatch,
`src/vt/vulkan/vulkan_context.cpp:174`) in `docs/ENVIRONMENT.md`.
5. `scripts/check-oracle-pins.py` — the sync `4f11dfc10` (2026-09-26) advanced
the parity pin to `a7c23ac96d` / `0.3.0.dev267` but left the restated
`<!--pin:commit-->`/`<!--pin:label-->` spans in `.agents/NOW.md`,
`docs/FEATURES.md`, `docs/benchmarks/how-we-measure.md`,
`docs/benchmarks/speculative-decoding.md` and
`docs/benchmarks/vllm-online-serving.md` at the prior pin. Restate the
current pin; the spans are the checker's only read of these surfaces.

## Gates

All five checkers exit 0; `check-agent-record.py` stays 0; the record suite
passes.

## Stop conditions

A checker whose expectation contradicts the tree's true state is reconciled to
the true state and the conflict named in the commit body, never silenced.

## Owed

- ISSUE-LOCAL-01M3WMKTBEDAHF226YSATSR7BF — the five-gate preflight repair this spec owns until it lands
16 changes: 16 additions & 0 deletions docs/BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,22 @@ registry-bound list in [FEATURES.md](FEATURES.md).
| Benchmark ID | Disposition | Detail |
|---|---|---|
| `qwen38-tensorfold-gap` | `BLOCKED_MISSING_ARTIFACTS`; no number | [Qwen3.8 TensorFold gap](benchmarks/qwen38-tensorfold-gap.md) |
| `at-a-glance` | portfolio snapshot, current W5/W6 state | [At a glance](benchmarks/at-a-glance.md) |
| `dwarfstar-gguf` | `DONE`: byte-exact decode, 1.144x vs DwarfStar `ds4` with `VT_V4_RESIDENT_W` | [DwarfStar, GGUF](benchmarks/dwarfstar-gguf.md) |
| `how-we-measure` | methodology reference for every binding number | [How we measure](benchmarks/how-we-measure.md) |
| `llama-cpp-cpu` | `DONE`: prefill 1.18x ahead, decode tie, memory parity (aarch64) | [llama.cpp, CPU](benchmarks/llama-cpp-cpu.md) |
| `memory` | `PASS`: peak PSS/RSS/GPU-memory ratios vs vLLM, 27B NVFP4 GB10 | [Memory](benchmarks/memory.md) |
| `mlx-lm-apple-m4` | INDICATIVE: 97.6% warm total vs MLX-LM, Qwen3-0.6B | [MLX-LM, Apple M4](benchmarks/mlx-lm-apple-m4.md) |
| `open-gaps` | ledger of open performance gaps and their next gates | [Open gaps](benchmarks/open-gaps.md) |
| `qwen38-27b-exl3-gb10` | EXL3 3.5bpw decode, with and without the DFlash2 draft | [Qwen3.8-27B EXL3 GB10](benchmarks/qwen38-27b-exl3-gb10.md) |
| `qwen38-27b-exl3-variadic-gb10` | EXL3 under a mixed-length serving load | [Qwen3.8-27B EXL3 variadic GB10](benchmarks/qwen38-27b-exl3-variadic-gb10.md) |
| `qwen38-27b-q4km-gfx1151` | Q4_K_M three-engine comparison on Strix Halo | [Qwen3.8-27B Q4_K_M gfx1151](benchmarks/qwen38-27b-q4km-gfx1151.md) |
| `reproduce` | entry points and recipes to reproduce each benchmark | [Reproduce](benchmarks/reproduce.md) |
| `speculative-decoding` | MTP/DFlash/n-gram/DSpark spec-decode rows vs vLLM | [Speculative decoding](benchmarks/speculative-decoding.md) |
| `tt-capture-default-decode` | Tenstorrent capture-default decode rate | [Tenstorrent capture-default decode](benchmarks/tt-capture-default-decode.md) |
| `tt-keepquant-27b-decode` | Tenstorrent keep-quant 27B decode, capture-default | [Tenstorrent keep-quant 27B decode](benchmarks/tt-keepquant-27b-decode.md) |
| `variadic-load-methodology` | methodology of the mixed-load serving benchmark | [Variadic-load methodology](benchmarks/variadic-load-methodology.md) |
| `vllm-online-serving` | the binding vLLM online-serving comparison grids | [vLLM, online serving](benchmarks/vllm-online-serving.md) |

## vLLM, online serving

Expand Down
Loading
Loading