From 31eb5e5431511c75974d0c152de8784ed5c7fd3c Mon Sep 17 00:00:00 2001 From: Yoav Date: Sun, 27 Sep 2026 19:51:30 -0400 Subject: [PATCH] fix(ORACLE-PIN-SURFACES): restate the a7c23ac96d pin, and re-anchor every claim the substitution made false The pin advanced in 4f11dfc10 (#3320) from e126687a9a to a7c23ac96d and the parity-pin authority moved with it. check-oracle-pins went red on eleven sites and every one was the same defect: the oracle record copy and the five prose surfaces in PIN_SURFACES still carried the old revision. Moving the seven pin-bearing values turns that gate green and leaves the tree asserting what the advance did not do, so this change moves the surrounding claims too rather than only the spans: - NOW.md claimed a gate HAD run at the pin and PASSED. It ran at e126687a9a on 2026-09-04, and #2817 is that advance's issue. It now records NO gate at the current pin, which is what FEATURES.md has been quoting it as saying. - oracles/vllm.md: the whole of "What this pin establishes" is measured at e126687a9a, so it is scoped to that pin explicitly; the device table's dgx row is relabelled the prior pin, the strix rows two pins back, the "exactly one row AT THE CURRENT PIN" paragraph is corrected to no rows (the 2026-09-22 advance added a pin and no measurement), and the residual's next row moves to the current pin. - The `oracle-pin` block's `gateable` field is now `no`, naming #3346. It read `yes`, carried over from the pin that earned it, which made the STRUCTURED field assert a measurement that does not exist at this revision. The September 22 sync report it pointed at records this tree's own C++ -- the `vllm` library build and `test_scheduler` 48/48 -- not an oracle build or a model run, and the one PORT-NOW commit in the range was ported from source reading rather than from a capture. AGENTS.md requires `no` until the oracle demonstrably builds and runs a model, and every other ungateable oracle in `.agents/oracles/` already says so. The prose caveat this change previously relied on cannot enforce it: check-oracle-pins.py:314-327 accepts a `yes` whose `evidence` path merely EXISTS in the tree, so pointing that field at a sync report left the checker reading `yes` on an unmeasured pin. The device table's rows are unaffected and still outrank the file-level field for their own devices. - The three benchmark pages stop dating the advance 2026-09-03 and citing #2817 for it, and stop claiming FlashInfer "moves to 0.6.18 at the new pin". The parity-pin block still reads 0.6.15.post1 and the 0.6.18 leg is refused by online_gate.py, so no published ratio has a 0.6.18 denominator. - how-we-measure.md's "no gate has run at the new pin" was stale in the other direction: #2794 closed COMPLETED on 2026-09-06 with TOKENGATE_VERDICT PASS at the prior pin. Four of five goldens remain owed and are unanchored. - .agents/NOW.md's rewritten Current gate paragraph had grown to 104 lines, past the 100-line budget check-now-current.py enforces. It is condensed to the live position (96 lines): the job ID, the verdict fields and the per-item evidence it drops already live in oracles/vllm.md's NOT-established items, which the paragraph now points at instead of duplicating. check-oracle-pins, check-device-leakage, check-now-current and the rest of the doc gates are green. check-env-doc is red at this base, and was red at the base this branch was cut from: 891007a6e (2026-08-29) added VT_VK_DISABLE, VT_VK_DISABLE_PAGED_ATTN and VT_VK_FENCE_TIMEOUT_MS to src/vt/vulkan/vulkan_ops.cpp without documenting them, so the gate has been red on main since that commit landed -- upstream's own scheduled CI at fce36733b fails on the same three names -- and the repair belongs to the Vulkan row, not to this change. check-agent-record still reports the nine _intake archive-evidence errors, which are the CRLF checkout artefact and are fixed by row/ENG-EOL-BYTE-EXACT, not by this change. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:codebuff/buffy [freebuff] --- .agents/NOW.md | 32 ++++---- .../ISSUE-LOCAL-01M3JM3HMK2F41A10P7T2WQXD1.md | 33 ++++++++ .agents/oracles/vllm.md | 82 +++++++++++++------ docs/FEATURES.md | 23 +++--- docs/benchmarks/how-we-measure.md | 29 +++++-- docs/benchmarks/speculative-decoding.md | 15 ++-- docs/benchmarks/vllm-online-serving.md | 10 ++- 7 files changed, 155 insertions(+), 69 deletions(-) create mode 100644 .agents/issues/ENG-RECORD-CLAIM-AGREEMENT/ISSUE-LOCAL-01M3JM3HMK2F41A10P7T2WQXD1.md diff --git a/.agents/NOW.md b/.agents/NOW.md index f9a4f16dee..23ec5a1ecd 100644 --- a/.agents/NOW.md +++ b/.agents/NOW.md @@ -22,24 +22,20 @@ no per-row change needs to touch this file at all. Token-exact (or ratified distributional) vs pinned vLLM; ≥ throughput and ≤ latency/memory on every axis, both gate models, reproduced 2–3x idle. See -[verification](verification.md). Pin: vLLM `e126687a9a` (0.28.1rc1.dev132) since -2026-09-03 (#2817). **A gate HAS now run at it and it PASSED** (2026-09-04, job -`7386f034-246a-4af5-9a04-f98aafffce54`, `dgx:gpu0`, 2h15m): the OPT candidate -captured at the target is byte-identical to the committed bar -- -`IDS mismatched_positions 0 of 96`, `IDS_BYTE_EQUAL True`, -`SELECTOR K=5 multi_valued_cells 0`, `TOKENGATE_VERDICT PASS`. The default -FLASH_ATTN backend produced the tokens, so the FA-on-GB10 risk did not fire. Our -arm's 96/96 carries over unchanged because the candidate's bytes are identical to -the bar it already passed. The BENCHMARK baselines are still measured at -`555967922` and still owe step 6 (#2818, `OPEN` at a 2026-09-12 read), **and the -other four strict goldens -- 27B W4A4, 32B-NVFP4A16, 35B, Coder -- are still -owed at the target.** The gate that passed is the OPT-125m one; it discharges -nothing about those four. They were anchored on #2794, which closed `COMPLETED` -on 2026-09-06, so the anchor is gone and not the obligation: it lives under -`## Owed` in [the tokengate spec](specs/upstream-sync-headpin-tokengate.md#L569), -unanchored by design, and whoever takes it files the issue then. Also at -[oracles/vllm.md](oracles/vllm.md#L93). - +[verification](verification.md). Pin: vLLM `a7c23ac96d` (0.3.0.dev267) since +2026-09-22 ([#3320](https://github.com/mudler/vllm.cpp/pull/3320)). **NO gate has run at it** -- the advance carries a sync +report, no oracle build, no capture, no measurement; its two recorded gates +(`vllm` library build, `test_scheduler` 48/48) are this tree's own C++. +**The gate that HAS run and PASSED ran at the PRIOR pin `e126687a9a`** +(2026-09-03, #2817): OPT-125m token-exact 96/96, `IDS_BYTE_EQUAL True`, +`TOKENGATE_VERDICT PASS` on 2026-09-04, so the candidate's bytes carry that +PASS over unchanged. The BENCHMARK baselines still owe step 6 (#2818) at +`555967922`, **and the other four strict goldens -- 27B W4A4, 32B-NVFP4A16, +35B, Coder -- are still owed at the target**; the OPT-125m PASS discharges +none of them. The obligation is unanchored by design (#2794 closed +`COMPLETED` 2026-09-06), lives under `## Owed` in [the tokengate +spec](specs/upstream-sync-headpin-tokengate.md); job IDs, verdict fields and +per-item evidence: [oracles/vllm.md](oracles/vllm.md). ## Next actions diff --git a/.agents/issues/ENG-RECORD-CLAIM-AGREEMENT/ISSUE-LOCAL-01M3JM3HMK2F41A10P7T2WQXD1.md b/.agents/issues/ENG-RECORD-CLAIM-AGREEMENT/ISSUE-LOCAL-01M3JM3HMK2F41A10P7T2WQXD1.md new file mode 100644 index 0000000000..cd5e558692 --- /dev/null +++ b/.agents/issues/ENG-RECORD-CLAIM-AGREEMENT/ISSUE-LOCAL-01M3JM3HMK2F41A10P7T2WQXD1.md @@ -0,0 +1,33 @@ +ID: ISSUE-LOCAL-01M3JM3HMK2F41A10P7T2WQXD1 +Title: the a7c23ac96d parity pin advanced in 4f11dfc10 but six restating surfaces kept e126687a9a, and the gate reads only the span +Row: ENG-RECORD-CLAIM-AGREEMENT +State: OPEN +Kind: bug +GitHub: - +Mirror: PENDING +Availability: FULL +Created: 2026-09-27 +Updated: 2026-09-27 +Closed: - + +## Problem + +The vLLM parity pin advanced in 4f11dfc10 (PR #3320, merged 2026-09-26) from e126687a9a828d513c01a07cd69f025f27d63280 to a7c23ac96d7806e7c7e7d862eadbce5a33529b94, and the authority block in .agents/upstream-sync.md moved with it. Nothing else did. check-oracle-pins went red with 11 errors and every one of them was the SAME defect at a different site: the oracle record copy and the five prose surfaces named in PIN_SURFACES still carried the old revision. + +Updating the seven pin-bearing values turns the gate green and leaves the tree asserting things that are now false, which is worse than the red. Concretely, after the spans move: + +1. .agents/NOW.md:25 reads "Pin: vLLM a7c23ac96d (0.3.0.dev267) since 2026-09-03 (#2817). A gate HAS now run at it and it PASSED (2026-09-04, job 7386f034..., dgx:gpu0)". No gate has run at a7c23ac96d. The capture was taken at e126687a9a on 2026-09-04 and #2817 is the issue for THAT advance, not this one. The date, the issue reference and the "at it" all bind to the wrong advance once the pin in the sentence moves. + +2. .agents/oracles/vllm.md, the whole of "## What this pin establishes, and what it does NOT", is written about e126687a9a: the thor build block names e126687a9a828d51... explicitly, item 2 says the declared token-exact gate at "this pin" was captured 2026-09-04, and item 3 owes step 6 at that pin. None of it was re-measured at a7c23ac96d. The 2026-09-22 sync report for this advance records exactly two gates, test_scheduler 48/48 and the vllm library build, and both are OUR tree, not the oracle. So the section keeps its heading and its contents while the pin above it now names a revision for which it establishes nothing. + +3. The dgx:gpu0 row of the device-scoped table still reads "**e126687a9a, the CURRENT pin**". It is not the current pin; it is the prior one. The two strix rows read "**5559679229, the PRIOR pin**", which after this advance is two pins back, and the residual paragraph under the table says "Whoever adds the next row at e126687a9a narrows it further" when the next row belongs at a7c23ac96d. + +4. docs/FEATURES.md:14 says "the parity pin since 2026-09-03" and :45 says the current pin "advanced on 2026-09-03". Both dates describe the e126687 advance. The same paragraph cites NOW.md as recording "**NO gate has run at it**" - a sentence NOW.md no longer contains, and which becomes TRUE again only once the "at it" in NOW.md is re-pointed at the pin that actually has no gate. + +5. docs/benchmarks/how-we-measure.md:21, speculative-decoding.md:4 and vllm-online-serving.md:15 all date the advance 2026-09-03 and cite #2817, and all three describe FlashInfer as moving to 0.6.18 "at the new pin". The authority block still records flashinfer_version = 0.6.15.post1 after the advance, and .agents/oracles/vllm.md item 3 records the 0.6.15.post1 to 0.6.18 step as NOT discharged. So the substitution would attach an unmeasured dependency to a pin nobody has built. + +The class is the one this row exists for: a fact corrected in one place and left standing in another, and here the correcting place is the one the gate reads, so the gate goes green over a tree that says the pin moved and, everywhere else, that a gate ran at it. Fixing only the spans is the change that makes the defect invisible. + +## Resolution + +- diff --git a/.agents/oracles/vllm.md b/.agents/oracles/vllm.md index 57cb975be4..8d62fada83 100644 --- a/.agents/oracles/vllm.md +++ b/.agents/oracles/vllm.md @@ -16,22 +16,45 @@ id = vllm role = primary upstream = https://github.com/vllm-project/vllm scope = every behavior vLLM implements — defaults, modes, errors, edge cases, and both correctness and speed gates -pin = e126687a9a828d513c01a07cd69f025f27d63280 -pin_label = 0.28.1rc1.dev132 -pinned_on = 2026-09-03 -gateable = yes -evidence = .agents/sync/2026-09-03-e126687-runhalf.md +pin = a7c23ac96d7806e7c7e7d862eadbce5a33529b94 +pin_label = 0.3.0.dev267 +pinned_on = 2026-09-22 +gateable = no +evidence = #3346 ``` ## What this pin establishes, and what it does NOT -**`gateable = yes` above says the oracle builds and runs. It does NOT say any -gate in this tree has been run against it.** Those are different statements and +**`gateable = no` above says the oracle has NOT been built and run at this pin. +It does NOT say any gate in this tree has been run against it.** Those are +different statements and this section keeps them apart, because the pin advanced ([#2817](https://github.com/mudler/vllm.cpp/issues/2817)) on a developer ruling that put step 6 after step 7, so the advance carries obligations it has not discharged. +**EVERY MEASUREMENT IN THIS SECTION IS AT `e126687a9a`, THE PRIOR PIN, AND NONE +OF IT WAS RE-MEASURED AT `a7c23ac96d`.** The pin moved a second time on +2026-09-22 ([#3320](https://github.com/mudler/vllm.cpp/pull/3320), `4f11dfc10`), +to unblock `MODEL-JEV`, and 1187 commits separate the two revisions, 232 of +them touching ported subtrees. The sync report for that advance +([`../sync/2026-09-22-a7c23ac96d.md`](../sync/2026-09-22-a7c23ac96d.md)) records +two gates and **both are this tree's own C++** -- `test_scheduler` 48/48 and the +`vllm` library build -- not the oracle, and the one PORT-NOW commit in the range +(`c64b15cde5`, the scheduler) was ported from source reading rather than from a +capture. **`gateable` is therefore `no`, naming [#3346](https://github.com/mudler/vllm.cpp/issues/3346) +as the issue that owes the build-and-run.** It used to read `yes` here, carried +over from the pin that did earn it, which made the structured field assert a +measurement that does not exist at this revision; `AGENTS.md` §"Pin vLLM" requires +`no` until the oracle demonstrably builds and runs a model, and every other +ungateable oracle in `.agents/oracles/` already says so. Note that +`scripts/check-oracle-pins.py:314-327` accepts a `yes` whose `evidence` path +merely EXISTS in the tree, so pointing that field at a sync report left the +checker reading `yes` on an unmeasured pin; the prose caveat could not enforce +what the field claimed. **Read every `this pin`, `this +revision` and `this one` below as `e126687a9a`**, and read the section as a record +of the prior pin that the current one has not inherited. + ### Established, on two boards **This heading used to read "and this is the whole of it", and it named only the @@ -156,30 +179,37 @@ measurement and its explicit non-claims are in [`../sync/2026-09-03-e126687-runhalf.md`](../sync/2026-09-03-e126687-runhalf.md) §6 and §7. -### The prior pin, for anyone reading a number taken under it +### The two pins before it, for anyone reading a number taken under one `5559679229bc961848b121ccdeaa8fa5d79bec98`, `0.26.0.dev0`, pinned 2026-07-26, FlashInfer `0.6.15.post1`, CUTLASS DSL `4.6.0`, transformers `5.14.1`. Every binding ratio published in `docs/benchmarks/` was measured against it. That -advance re-captured goldens on its own oracle and recorded zero real drift; this -one has re-captured ONE of five, which is the difference §"NOT established" -item 2 is about. +advance re-captured goldens on its own oracle and recorded zero real drift; the +advance to `e126687a9a` re-captured ONE of five, which is the difference +§"NOT established" item 2 is about; and the advance to `a7c23ac96d` re-captured +NONE of them, because nothing was captured against it. `e126687a9a828d513c01a07cd69f025f27d63280`, +`0.28.1rc1.dev132`, pinned 2026-09-03, is the pin between the two, and the whole +of §"What this pin establishes" above is measured there. ## Device-scoped gateability -`gateable = yes` above is a property of the oracle, not a promise about every +**A row below is a DEVICE measurement and outranks the file-level +`gateable = no` for that device.** `gateable` in the block above is a +whole-oracle disposition, not a promise about every board. Where a device has been MEASURED to build and run a gate model, it is recorded here with the evidence; absence from this table means unmeasured, never unsupported. One row per measurement, appended by the change that made it. -**AT THE CURRENT PIN THIS TABLE NOW HAS EXACTLY ONE ROW, and its narrowness is -the honest reading of it.** The `dgx:gpu0` row is the 2026-09-04 token gate: a -source build at `e126687a9a` served `facebook/opt-125m` and reproduced the -committed golden byte-for-byte. That answers "does a gate model run at -`e126687a9a`" on one device with one model, and it answers nothing about the -production checkpoints the binding rows use. **This paragraph used to say the -table was empty and that nothing in this file answered that question; the capture -falsified it.** +**AT THE CURRENT PIN THIS TABLE HAS NO ROWS AT ALL, and the `dgx:gpu0` row below +is at the PRIOR pin.** The `dgx:gpu0` row is the 2026-09-04 token gate: a source +build at `e126687a9a` served `facebook/opt-125m` and reproduced the committed +golden byte-for-byte. That answers "does a gate model run at `e126687a9a`" on one +device with one model, and it answers nothing about the production checkpoints the +binding rows use, nothing about `thor:gpu0`, and **nothing at all about +`a7c23ac96d`, which no board in this tree has built.** **This paragraph said the +table was empty, and then that it held exactly one row AT THE CURRENT PIN; the +first was falsified by the 2026-09-04 capture and the second by the 2026-09-22 +advance, which added a pin to the authority and no measurement to this table.** The two `strix:gpu0` rows were measured at the PRIOR pin `5559679229`: the first row's evidence file records `SETUPTOOLS_SCM_PRETEND_VERSION=0.26.0.dev0+g5559679229` @@ -215,9 +245,9 @@ its golden was captured under. | device | arch | pin measured at | measured | builds | runs a gate model | evidence | |---|---|---|---|---|---|---| -| `dgx:gpu0` | GB10, `sm_121a` (CUDA 13.0, driver 580.173.02) | **`e126687a9a`, the CURRENT pin** | 2026-09-04 | yes — from source, `PORCELAIN_LINES=0`, wheel `vllm-0.28.1rc1.dev132+ge126687a9-…-linux_aarch64.whl` | yes — `facebook/opt-125m` bf16, 6 prompts x 16 greedy tokens, `--runs 5`: `IDS mismatched_positions 0 of 96`, `IDS_BYTE_EQUAL True`, `TOKENGATE_VERDICT PASS`. The narrowest gate model in the tree; says nothing about the production checkpoints | [`opt125m-token-gate-e126687-dgx-20260904.md`](../../docs/bench-evidence/opt125m-token-gate-e126687-dgx-20260904.md) | -| `strix:gpu0` | `gfx1151` (RDNA 3.5, Radeon 8060S, ROCm 7.2.4) | **`5559679229`, the PRIOR pin** | 2026-09-03 | yes | yes — Qwen3.8-27B Q4_K_M GGUF, 6 prompts x 48 greedy tokens, reproducible | [`oracle-vllm-gfx1151-20260903.md`](../../docs/bench-evidence/oracle-vllm-gfx1151-20260903.md) | -| `strix:gpu0` | `gfx1151` (RDNA 3.5, Radeon 8060S, ROCm 7.2.4) | **`5559679229`, the PRIOR pin** | 2026-09-03 | yes | yes — and SCORED a gate: `prompt_logprobs` teacher-forcing over 6 x 48 steps, both configurations, negative control discriminating at 21.24 nats | [`q4km-neartie-vllm-oracle-20260903.md`](../../docs/bench-evidence/q4km-neartie-vllm-oracle-20260903.md) | +| `dgx:gpu0` | GB10, `sm_121a` (CUDA 13.0, driver 580.173.02) | **`e126687a9a`, the prior pin** | 2026-09-04 | yes — from source, `PORCELAIN_LINES=0`, wheel `vllm-0.28.1rc1.dev132+ge126687a9-…-linux_aarch64.whl` | yes — `facebook/opt-125m` bf16, 6 prompts x 16 greedy tokens, `--runs 5`: `IDS mismatched_positions 0 of 96`, `IDS_BYTE_EQUAL True`, `TOKENGATE_VERDICT PASS`. The narrowest gate model in the tree; says nothing about the production checkpoints | [`opt125m-token-gate-e126687-dgx-20260904.md`](../../docs/bench-evidence/opt125m-token-gate-e126687-dgx-20260904.md) | +| `strix:gpu0` | `gfx1151` (RDNA 3.5, Radeon 8060S, ROCm 7.2.4) | **`5559679229`, two pins back** | 2026-09-03 | yes | yes — Qwen3.8-27B Q4_K_M GGUF, 6 prompts x 48 greedy tokens, reproducible | [`oracle-vllm-gfx1151-20260903.md`](../../docs/bench-evidence/oracle-vllm-gfx1151-20260903.md) | +| `strix:gpu0` | `gfx1151` (RDNA 3.5, Radeon 8060S, ROCm 7.2.4) | **`5559679229`, two pins back** | 2026-09-03 | yes | yes — and SCORED a gate: `prompt_logprobs` teacher-forcing over 6 x 48 steps, both configurations, negative control discriminating at 21.24 nats | [`q4km-neartie-vllm-oracle-20260903.md`](../../docs/bench-evidence/q4km-neartie-vllm-oracle-20260903.md) | **A residual this table exists to hold, and does not yet close.** Every checker reads the `gateable` field, not the prose above it, and no field names the device @@ -233,8 +263,10 @@ read 2026-09-12). The residual is therefore narrower and still open: the gap is closed for ONE board and ONE gate model, `facebook/opt-125m` at six prompts and sixteen tokens. It is not closed for `thor:gpu0`, for `strix:gpu0`, or for any production checkpoint at this pin, and no field records that limit. Whoever adds -the next row at `e126687a9a` narrows it further; nobody closes it by reading the -field. +the next row at `a7c23ac96d` narrows it further, and until one exists the +`gateable` in the block above is `no` for want of a build-and-run at this pin; a +device row does not lift that, it only ever narrows where a `no` is allowed to +apply. **gfx1151 needs five packages a bare ROCm image does not carry**, and each of their absences presents as a device failure rather than as a provisioning gap: diff --git a/docs/FEATURES.md b/docs/FEATURES.md index af3dc693e2..f8ae716b88 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -11,8 +11,8 @@ the agent-facing parity inventory with upstream file references see **Legend.** ✅ supported and gated. ◐ partial, usable with named gaps. ☐ not yet. n/a means the feature does not apply to that engine's design. -Reference versions: vLLM 0.28.1rc1.dev132 (`e126687a9a`, the parity pin since -2026-09-03), SGLang v0.5.15, llama.cpp `b10451`, MLX-LM as of 2026-07. Rows +Reference versions: vLLM 0.3.0.dev267 (`a7c23ac96d`, the parity pin since +2026-09-22, [#3320](https://github.com/mudler/vllm.cpp/pull/3320)), SGLang v0.5.15, llama.cpp `b10451`, MLX-LM as of 2026-07. Rows describing what vLLM has were read at the PRIOR pin `555967922` unless they say otherwise; the 290-commit-range PORT-NOW queue for the advance is classified and unworked (#2611). Competitor columns describe what those projects ship, and @@ -41,14 +41,17 @@ can ask `git merge-base --is-ancestor` rather than assume the result survived. The `GlmMoeDsaForCausalLM` ROCm arm is the worked example: the run was real at `9f3e6e223`, and #2511 later withdrew the premise it depended on. -*Results.* Every "vs vLLM" figure on this page was captured at the **prior** -parity pin `555967922` and **has not been re-validated** at the -current pin `e126687a9a`, which advanced on -2026-09-03. `.agents/oracles/vllm.md` states in its own words that the pin -advance "does NOT say any gate in this tree has been run against it", and -`.agents/NOW.md` records "**NO gate has run at it**". The rows below name -`vLLM 0.25.0` where that is the version they were measured against. Tracked -by #2794 (goldens predate the pin) and #2817 (the advance). +*Results.* Every "vs vLLM" figure on this page was captured at parity pin +`555967922`, two advances back, and **has not been re-validated** at the +current pin `a7c23ac96d`, which advanced on +2026-09-22 from `e126687a9a` ([#3320](https://github.com/mudler/vllm.cpp/pull/3320)). `.agents/oracles/vllm.md` states in its own words +that the pin advance "does NOT say any gate in this tree has been run against +it", and `.agents/NOW.md` records "**NO gate has run at it**" -- true of the +current pin, and true of the whole of §"What this pin establishes" in that file, +which is measured at `e126687a9a`. One token gate DID run and pass, at that +prior pin, on 2026-09-04. The rows below name `vLLM 0.25.0` where that is the +version they were measured against. Tracked by #2794 (goldens predate the pin), +#2817 (the 2026-09-03 advance) and [#3320](https://github.com/mudler/vllm.cpp/pull/3320) (the 2026-09-22 one). ## At a glance diff --git a/docs/benchmarks/how-we-measure.md b/docs/benchmarks/how-we-measure.md index 6bf7696dd8..19fa98ce6c 100644 --- a/docs/benchmarks/how-we-measure.md +++ b/docs/benchmarks/how-we-measure.md @@ -18,11 +18,19 @@ memory compete; end-to-end wall-clock on a cold page cache is unusable there, and steady-state per-step timing or `nsys` GPU-busy is the anchor. The 2026-08-06 #77-slip tree-revert changed no benchmark content or number. -**Oracle pin.** vLLM 0.28.1rc1.dev132 (`e126687a9a`) since 2026-09-03, with -FlashInfer `0.6.18` and CUTLASS DSL `4.6.2`. +**Oracle pin.** vLLM 0.3.0.dev267 (`a7c23ac96d`) since 2026-09-22 ([#3320](https://github.com/mudler/vllm.cpp/pull/3320)), +and **no oracle has been built at it.** FlashInfer `0.6.18` and CUTLASS DSL +`4.6.2` are the versions the `e126687a9a` oracle was INSTALLED with +([`sync/2026-09-02-e126687.md`](../../.agents/sync/2026-09-02-e126687.md) §5.2) +and the ones step 6 owes, not anything measured at the current pin: the +`parity-pin` block still reads `flashinfer_version = 0.6.15.post1` and +`tools/bench/online_gate.py` refuses a leg that does not match it +([`sync/2026-09-03-e126687-step6.md`](../../.agents/sync/2026-09-03-e126687-step6.md), +`REFUSED_AT_FLASHINFER`), so no published ratio on these pages ever ran against a +`0.6.18` oracle ([#2818](https://github.com/mudler/vllm.cpp/issues/2818)). -**EVERY BINDING RATIO ON THESE PAGES WAS MEASURED AGAINST THE PREVIOUS PIN**, -vLLM 0.26.0.dev0 (`55596792`) plus transformers 5.14.1, built from source for +**EVERY BINDING RATIO ON THESE PAGES WAS MEASURED AGAINST A PIN TWO ADVANCES +BACK**, vLLM 0.26.0.dev0 (`55596792`) plus transformers 5.14.1, built from source for sm_121a, the running oracle reporting `0.23.1rc1.dev1511+g555967922` with FlashInfer `0.6.15.post1`, selected by explicit path and asserted per leg. The pin advanced before those baselines were re-measured, by a developer ruling that @@ -31,10 +39,15 @@ committed harness refuses to run at any revision but the pinned one. **The re-measurement is owed against the new pin ([#2818](https://github.com/mudler/vllm.cpp/issues/2818)), it covers five rows across `vllm-online-serving` and `speculative-decoding`, and a red there requires -reverting the pin rather than re-arguing the rows.** No gate of any kind has yet -run at the new pin: the declared token-exact gate is owed by -[#2794](https://github.com/mudler/vllm.cpp/issues/2794) and every committed -golden predates the advance. Speed figures labelled 0.25.0 ran the ROLLBACK the +reverting the pin rather than re-arguing the rows.** No gate of any kind has +run at the current pin, and the sentence this one replaced was itself stale: it +claimed nothing had run "at the new pin" while the declared token-exact gate +owed by [#2794](https://github.com/mudler/vllm.cpp/issues/2794) HAD run and +PASSED, at `e126687a9a`, on 2026-09-04, `TOKENGATE_VERDICT PASS`, and #2794 +closed `COMPLETED` on 2026-09-06. What that capture closed is ONE of five +goldens; the other four -- 27B W4A4, 32B-NVFP4A16, 35B, Coder -- are still owed +and are no longer anchored on an issue (`.agents/oracles/vllm.md` §"NOT +established" item 2). Speed figures labelled 0.25.0 ran the ROLLBACK the harness enforced until 2026-08-12 and are SUPERSEDED, never binding (#520). Correctness re-validated bit-identical across the advance, zero golden drift. The llama.cpp oracle is stock `b10451` since 2026-08-16, `gateable = no` until diff --git a/docs/benchmarks/speculative-decoding.md b/docs/benchmarks/speculative-decoding.md index 985ad2317e..1173cee0c5 100644 --- a/docs/benchmarks/speculative-decoding.md +++ b/docs/benchmarks/speculative-decoding.md @@ -1,11 +1,16 @@ # Speculative decoding -**THE PIN MOVED UNDER TWO OF THESE ROWS.** The parity pin -advanced to `e126687a9a` on 2026-09-03 +**THE PIN MOVED UNDER TWO OF THESE ROWS, TWICE.** The parity pin +advanced to `a7c23ac96d` on 2026-09-22 +([#3320](https://github.com/mudler/vllm.cpp/pull/3320)), from `e126687a9a`, which it reached on 2026-09-03 ([#2817](https://github.com/mudler/vllm.cpp/issues/2817)). The **MTP** row and -the **DFlash** row carry a vLLM denominator measured against the PREVIOUS pin -`555967922` with FlashInfer `0.6.15.post1`, which is the oracle's attention -backend on that path and moves to `0.6.18` at the new pin. Both owe a +the **DFlash** row carry a vLLM denominator measured against a pin TWO advances +back, `555967922`, with FlashInfer `0.6.15.post1`, which is the oracle's attention +backend on that path. **FlashInfer is `0.6.18` on the oracle this tree has +actually built and on every pin since, and that step is STILL NOT discharged** -- +the `parity-pin` block reads `0.6.15.post1` and the `0.6.18` leg is REFUSED by +the harness, so neither row's denominator has ever run against a `0.6.18` +oracle. Both owe a re-measurement ([#2818](https://github.com/mudler/vllm.cpp/issues/2818)), and the DFlash re-take must record the oracle's SELECTED backend rather than assume it carries over. The advance preceded the re-measurement by a developer ruling that diff --git a/docs/benchmarks/vllm-online-serving.md b/docs/benchmarks/vllm-online-serving.md index 40dc8e712d..2ee74e8ef2 100644 --- a/docs/benchmarks/vllm-online-serving.md +++ b/docs/benchmarks/vllm-online-serving.md @@ -12,10 +12,14 @@ The first series free of both, at the pin, graphed, and at a pinned clock is in [the benchmark record](../../.agents/benchmark-record.md). **THE PIN MOVED UNDER THESE ROWS, AND THEY HAVE NOT BEEN RE-MEASURED.** The -parity pin advanced to `e126687a9a` on 2026-09-03 +parity pin advanced to `a7c23ac96d` on 2026-09-22 +([#3320](https://github.com/mudler/vllm.cpp/pull/3320)), from `e126687a9a`, which it reached on 2026-09-03 ([#2817](https://github.com/mudler/vllm.cpp/issues/2817)). Every row below that -says "at the pin" was measured against the PREVIOUS pin `555967922`, with -FlashInfer `0.6.15.post1`. FlashInfer moves to `0.6.18` at the new pin and is on +says "at the pin" was measured against a pin TWO advances back, `555967922`, with +FlashInfer `0.6.15.post1`. **FlashInfer is `0.6.18` on the oracle this tree has +actually built and on every pin since, and the step is STILL NOT discharged** -- +the `parity-pin` block reads `0.6.15.post1` and the `0.6.18` leg is REFUSED by +`online_gate.py`, so no row here has a `0.6.18` denominator. It is on the executed path of these numbers on both sides — it is the NVFP4 GEMM under the denominator and the CUTLASS source tree our own arm compiles against. **Four rows here owe a re-measurement**: the 27B NVFP4 `nvidia` and 35B-A3B NVFP4 rows, the