Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/backend-matrix.md

Large diffs are not rendered by default.

66 changes: 63 additions & 3 deletions .agents/environment.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,56 @@ Three consequences for anyone sizing work here:
at 73.9 GiB `VmHWM`; the host side is now 31 GiB total. Run a CPU comparison
on `thor` or `dgx`, or on-box against an oracle instead.

### `garlic-clove` — Intel Arc Pro B60

**It is NOT in the fleet table above, and it is not a fleet device.** Do not
read its absence as unavailability: it is reachable, and it is the only Intel
GPU on this estate. It is deliberately absent from the `rc` table because it is
not enrolled in resource-controller — `rc devices` could not be consulted from
the shell this was written in (no `rc` client), and nothing here may claim
membership it has not verified. **Therefore the lease rule does not apply, and
the file mutex does: take `${GPU_LOCK:-$HOME/gpu.lock}` ON THE BOX, as the
non-fleet-device clause of §"Reaching a GPU" requires.** Do not `ssh` in and
start GPU work unguarded, and do not add it to the fleet table until `rc
devices` actually reports it.

| | |
|---|---|
| Host | `garlic-clove`, Tailscale name, `100.67.232.8`, Linux, account `yoav` (lowercase) |
| GPU | **Intel Arc Pro B60 Graphics (BMG G21)**, `8086:e211`, ASUS subsys `1849:6023`, `xe` kernel driver |
| Vulkan | device API **1.4.354**, conformance 1.4.0.0, Mesa **26.2.3** (kisak PPA), LLVM 21.1.8; `VK_KHR_cooperative_matrix` = true, `VK_KHR_shader_bfloat16`, `VK_KHR_shader_integer_dot_product` |
| Device type | **`PHYSICAL_DEVICE_TYPE_DISCRETE_GPU`** |
| Memory | 20.91 GiB `DEVICE_LOCAL` heap + 23.44 GiB host heap; `memoryTypes[3]`/`[6]` = `DEVICE_LOCAL \| HOST_VISIBLE \| HOST_COHERENT` (0x0007) via ReBAR |
| Host | Ubuntu 24.04.4 LTS, kernel 7.0.0-34, i7-10700 8c/16t, 31 GiB RAM, 465 GiB NVMe (159 GiB free) |
| Toolchain | cmake 3.28.3, gcc 13.3.0, ninja 1.11.1, git, py3.12. **No `icpx`, no `sycl-ls`, no `/dev/accel/*`.** |
| Assets | **no model weights** — `~/.cache/huggingface` is 3.5 MB; only vocab-only GGUFs under `~/llama-maple/models` |
| Checkout | `~/vllm.cpp`, was on `row/BACKEND-VULKAN-TQ1_0-finish` with `VLLM_CPP_VULKAN=ON`, `BUILD EXIT 0`; that work appears to have landed on `main` as #2248, so the branch needs a fetch/prune, not a merge |

**Its load-bearing property is ReBAR, and that is not a property of Arc — but it
is a performance property, not a correctness one.** ReBAR is what makes the
allocation *device-local AND* host-addressable on a *discrete* card, which is
what the GGUF keep-quant CPU fall-through and the portable reference tier both
want. `VulkanContext` (`vulkan_context.cpp:872-875`) does not require
`DEVICE_LOCAL`: it prefers the device-local host-visible type and falls back to
plain `HOST_VISIBLE | HOST_COHERENT`, and `AllocBuffer` maps whichever it picks,
so `DeviceMemoryIsHostAddressable()` stays true without ReBAR. A discrete board
without ReBAR therefore runs those paths **slower**, with the GPU reaching the
allocation across the bus, not incorrectly. Host visibility is the correctness
requirement; device locality is the performance preference. See
[`specs/vulkan-full-support.md`](specs/vulkan-full-support.md) §4 and the
comment at `src/vt/vulkan/vulkan_ops.cpp` in the `BACKEND-VULKAN-KEEPQUANT`
block.

**Two cosmetic warts, recorded because a clean enumeration is a gate here.** A
stale `dzn_icd.json` (declaring `api_version 1.1.354`) makes the loader print
`Received return code -9 from call to vkCreateInstance in ICD libvulkan_dzn.so.
Skipping this driver` on every enumeration. It is HARMLESS — the working
`intel_icd.json` provides the device, and `vulkaninfo --summary` lists the B60
as GPU0 — but it will pollute any recorded enumeration. The loader's *instance*
version is also 1.3.275 while the *device* is 1.4.354; the backend `dlopen`s the
loader and only the device capability matters, but the two numbers should not
be quoted interchangeably.

### `orin:gpu0` needs L4T CUDA 12.6, and the DGX recipe breaks it

**Do not install `cuda-toolkit-13-*` from the generic `sbsa` repo on `orin`.** That
Expand Down Expand Up @@ -2340,6 +2390,16 @@ inner 4096, state 128; context 262144.
Enumeration and the clean-tree rule:
[`specs/oracle-llamacpp-repin-stock.md`](specs/oracle-llamacpp-repin-stock.md).

- **No Intel GPU exists on any box here**, so `BACKEND-XPU` end-to-end work is
HW-BLOCKED; only policy-port, compile coverage and oneAPI CPU-device unit
numerics are available.
- **An Intel GPU now exists on this estate: `garlic-clove`, an Intel Arc Pro B60.**
This line used to read "**No Intel GPU exists on any box here**, so
`BACKEND-XPU` end-to-end work is HW-BLOCKED", and that was true when written
but is FALSE as of 2026-09-27. The board is the hardware `VK-I` was scoped
against; see [garlic-clove](#garlic-clove--intel-arc-pro-b60) below. **`BACKEND-XPU`
nonetheless stays `SPIKE`/HW-blocked for a DIFFERENT and still-true reason: the
SYCL toolchain is absent, not the GPU.** Measured on the box 2026-09-27 — no
`icpx`, no `sycl-ls`, and no `/dev/accel/*` nodes, so there is no Level Zero
driver userspace to dispatch through. The Level Zero *runtime* libraries are
installed (`libze_loader.so.1`, `libze_intel_gpu.so.1`) and are not
sufficient. So the available work remains policy-port, compile coverage and
oneAPI CPU-device unit numerics, and closing `BACKEND-XPU` now needs an
oneAPI install rather than hardware acquisition.
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
ID: ISSUE-LOCAL-01M3JGPW1FN5SG506ANGBT3C0Q
Title: VK-I: B60 ReBAR retires the discrete staging path; the 'B60 is integrated' comment and the 'no Intel GPU' registry line are both false
Row: BACKEND-VULKAN
State: OPEN
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-27
Updated: 2026-10-02
Closed: -

## Problem

Three records about the Intel Arc Pro B60 are wrong, and one planned deliverable is answered by hardware rather than by code.

(1) FALSE CODE COMMENT. src/vt/vulkan/vulkan_ops.cpp:2109 asserts 'The B60 is integrated'. Measured on garlic-clove with vulkaninfo, the device reports deviceType = PHYSICAL_DEVICE_TYPE_DISCRETE_GPU, with two separate heaps (20.91 GiB DEVICE_LOCAL, 23.44 GiB host-visible). The CODE is right for the stated WRONG reason: VulkanBackend::DeviceMemoryIsHostAddressable() (vulkan_backend.cpp:135) returns true unconditionally, and vulkan_context.cpp:873 prefers a DEVICE_LOCAL|HOST_VISIBLE|HOST_COHERENT type, which memoryTypes[3] and memoryTypes[6] (propertyFlags 0x0007) satisfy. The property the code depends on holds; the stated reason does not. This is a trap rather than a typo: the comment grounds host addressability in the card being integrated, but the allocator guarantees it on any board -- vulkan_context.cpp:872-875 prefers DEVICE_LOCAL|HOST_VISIBLE|HOST_COHERENT and falls back to plain HOST_VISIBLE|HOST_COHERENT, refusing to initialize only if that fails too. A non-ReBAR Arc card is therefore still host-addressable through the second type; what it loses is device locality, so the keep-quant fall-through and the portable reference tier run slower there, not unsafe.

(2) FALSE REGISTRY LINE. .agents/environment.md asserts 'No Intel GPU exists on any box here, so BACKEND-XPU end-to-end work is HW-BLOCKED'. garlic-clove is an Intel Arc Pro B60 (8086:e211, ASUS subsys 1849:6023) on the xe driver.

(3) VK-I's STAGING-PATH DELIVERABLE IS ANSWERED, NOT BUILDABLE. vulkan-full-support.md scopes VK-I as 'The staging path for non-host-visible memory, and the gate re-run where Vulkan actually matters', and records the 2026-08-06 decision 'GB10 first, acquire later'. The acquire has happened, but the first half needs no code: ReBAR maps the B60's 20.91 GiB as host-visible, so the non-host-visible condition VK-I was written to handle does not arise on this card. Recording that as a measurement retires the risk; writing a staging path against it would be code for a condition this hardware does not have.

Also stale, and the reason this was not caught: the BACKEND-VULKAN matrix cell still describes the 2026-07-22 skeleton (8 native ops, 'no model runs on Vulkan'). The tree now registers 35 Vulkan ops including the full GDN/SSM set, GGUF keep-quant/TQ1_0, EXL3 and MoE, and vulkan_ops.cpp:2109 itself names this board.

## Resolution

-
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
ID: ISSUE-LOCAL-01M3JGXN1QTRE1D7WX3JMV0KE2
Title: There is no runtime device selector: VLLM_CPP_DEVICE is read nowhere, so every non-CUDA gate selects its device by accident
Row: BACKEND-VULKAN
State: OPEN
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-27
Updated: 2026-09-27
Closed: -

## Problem

The engine picks its device in src/vllm/entrypoints/model_loader.cpp via CurrentPlatform(), which walks {kCUDA, kXPU, kVULKAN, kMETAL, kCPU} and takes the FIRST backend that probed a device. Nothing overrides that. The env var the campaign recorded, VLLM_CPP_DEVICE, is read NOWHERE: grep over the whole tree returns only record files, never src/, include/, tests/ or examples/.

So `VLLM_CPP_DEVICE=vulkan test_opt_paged_engine` selected Vulkan only because those builds had no CUDA compiler. benchmark-record.md already MEASURED both ways on one source: with /usr/local/cuda/bin on PATH at configure time the identical command reports 'the engine selected device type 1' (kCUDA) and passes 6/6; without it, 'device type 3' (kVULKAN), 6/6, 0 declines.

Why this is filed now rather than read and skipped: the estate gained a SECOND accelerator host (garlic-clove, Intel Arc Pro B60, no CUDA toolchain at all). A gate that selects its device by build accident is exactly the failure that a new box invites -- on a host with no CUDA the placebo is indistinguishable from a working selector, so the first Vulkan-only venue cannot tell the difference. The test already prints the truth ('the engine selected device type N') and load-direct-upload.md already uses the correct idiom, so the defect is that no selector EXISTS, not that the evidence is unobtainable.

Scope: add a real runtime device override read at the SelectQueue seam, make it fail LOUDLY on an unknown value rather than silently falling through, and pin the selection in the test's own assertions. Then replace the remaining VLLM_CPP_DEVICE command citations in the record with the BACKEND PROOF form. Not attempted in the records-only VK-I change, because it is a code change to the device seam and needs its own row, spec and gate.

## Resolution

-
91 changes: 85 additions & 6 deletions .agents/specs/vulkan-full-support.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,11 @@ a win over *the Vulkan maturity floor on the box we own*, and the record must sa
exactly that rather than "we beat llama.cpp". The claim that would matter to a
user — Vulkan winning where CUDA/ROCm/Metal do not exist — is `VK-I`, and it is
hardware-blocked until an RDNA or Arc board is acquired (user decision 2026-08-06:
GB10 first, acquire later).
GB10 first, acquire later). **The Arc half of that acquisition HAPPENED: see
§4 and §6.2 for the board (`garlic-clove`, Intel Arc Pro B60) and for the
measurement that answers `VK-I`'s staging-path half without code. The RDNA arm
is still unacquired, and the gate re-run — the half that actually matters —
is still owed.**

---

Expand Down Expand Up @@ -248,7 +252,8 @@ our own CUDA paged kernel**, recorded as a partial-from-scratch entry in
|---|---|---|
| **GB10 on `dgx.casa`** | YES — `NVIDIA GB10`, `INTEGRATED_GPU`, Vulkan API 1.4.312, vendor `0x10de`, 249 device extensions incl. **`VK_KHR_cooperative_matrix` v2** and **`VK_NV_cooperative_matrix2`**, `VK_KHR_shader_float16_int8`, `VK_KHR_{8,16}bit_storage`, `VK_KHR_shader_integer_dot_product`, `VK_KHR_timeline_semaphore`, `VK_EXT_memory_budget`, `VK_KHR_buffer_device_address`. One 89.72 GiB `DEVICE_LOCAL` heap with a `DEVICE_LOCAL|HOST_VISIBLE` type — unified | **PRIMARY. Correctness oracle box AND the optimization target** (user decision 2026-08-06). Both llama.cpp coopmat tiers are reachable |
| **`llvmpipe` (dev box)** | YES — Vulkan 1.4.318, CPU, `mesa-vulkan-drivers` | GPU-free CI correctness. **Never a speed venue** |
| **AMD RDNA / Intel Arc** | **NO — none on any box** | `VK-I`. Deferred by user decision; the only venue where a Vulkan win means something to a user, and the only thing that exercises the missing staging path |
| **Intel Arc Pro B60 on `garlic-clove`** | **YES — acquired, measured 2026-09-27.** `Intel(R) Arc(tm) Pro B60 Graphics (BMG G21)`, `8086:e211`, ASUS subsys `1849:6023`, `xe` driver, `PHYSICAL_DEVICE_TYPE_DISCRETE_GPU`, device API **1.4.354** (conformance 1.4.0.0), Mesa 26.2.3 / LLVM 21.1.8, `VK_KHR_cooperative_matrix` = true. Heaps 20.91 GiB `DEVICE_LOCAL` + 23.44 GiB host. **NOT an `rc` fleet device** — file mutex, not a lease | `VK-I`, PARTIALLY ANSWERED — see §6.2. The "acquire later" half of the 2026-08-06 decision is DONE |
| ~~**AMD RDNA**~~ | **NO — still none on any box** | `VK-I`'s RDNA arm. A discrete AMD board would need ReBAR for the same reason, and is still unacquired |

**Premise update — the 2026-07-22 toolchain constraint is STALE.** That spec
determined the committed-SPIR-V route partly because *"neither box grants sudo"*
Expand Down Expand Up @@ -327,6 +332,71 @@ the umbrella, not a substitute for them.
| **VK-H** | **Attention variants + samplers** (16 ops) | B (samplers), G (attn variants) | **83/83 — closes the op surface** |
| **VK-I** | **AMD/RDNA (or Arc) bring-up** | hardware acquisition | The staging path for non-host-visible memory, and the gate re-run where Vulkan actually matters |

### 6.2 `VK-I` PARTIALLY ANSWERED — ReBAR retires the staging path — 2026-09-27

`VK-I` was the one sub-project blocked purely on acquisition, and §4 recorded
the decision as **"GB10 first, acquire later"** (2026-08-06). The board is now
on the estate: **`garlic-clove`, an Intel Arc Pro B60.** Its deliverable splits,
and the split is not the one the row was written against.

**THE STAGING-PATH HALF IS ANSWERED BY HARDWARE, NOT BY CODE, AND NO CODE
SHOULD BE WRITTEN FOR IT.** The deliverable was "the staging path for
**non-host-visible** memory". MEASURED on `garlic-clove` 2026-09-27 with
`vulkaninfo`: the device is `PHYSICAL_DEVICE_TYPE_DISCRETE_GPU` with two heaps
(20.91 GiB `DEVICE_LOCAL`, 23.44 GiB host), **but** `memoryTypes[3]` and
`memoryTypes[6]` expose `DEVICE_LOCAL | HOST_VISIBLE | HOST_COHERENT`
(propertyFlags `0x0007`) on the device-local heap. Resizable BAR maps the VRAM
into the host address space, so the condition `VK-I` was built to survive —
device-local memory the host may not dereference — **does not arise on this
card**. `VulkanContext`'s existing preference for a device-local host-visible
type (`vulkan_context.cpp:873`) already selects it, so
`VulkanBackend::DeviceMemoryIsHostAddressable()` is sound here without a
staging copy.

**THIS RETIRES A RISK; IT DOES NOT DELETE A REQUIREMENT.** The property is a
property of **ReBAR**, not of Arc and not of this backend — and it is a
**performance** property, not a correctness one. `VulkanContext`
(`vulkan_context.cpp:872-875`) does not require `DEVICE_LOCAL` at all: it
PREFERS `DEVICE_LOCAL | HOST_VISIBLE | HOST_COHERENT`, and when that lookup
fails it FALLS BACK to `HOST_VISIBLE | HOST_COHERENT` without `DEVICE_LOCAL`,
initializing only if that fails too. `AllocBuffer` allocates and
`vkMapMemory`-maps whichever type was selected, and
`DeviceMemoryIsHostAddressable()` returns true unconditionally
(`vulkan_backend.cpp:130-135`) because every allocation is mapped.

So a discrete card without ReBAR does **not** turn the GGUF keep-quant CPU
fall-through or the portable reference tier into corruption. Host visibility
is what correctness needs, the fallback supplies it, and the host vec_dot
kernel keeps reading memory it can address. What the card without ReBAR costs
is that the GPU reaches that memory across the bus rather than hitting VRAM
locally — **slowness, not unsafety**, and the staging path `VK-I` exists to
buy that back. So the standing requirement is the one already in the code:
keep *preferring* a device-local host-visible type and let the ordered
fallback carry correctness where none exists. `unified_memory_` records which
of the two happened, and it is the performance branch, not a safety one. That
is now written down at the `BACKEND-VULKAN-KEEPQUANT` block in
`src/vt/vulkan/vulkan_ops.cpp`, which until 2026-09-27 carried the false claim
**"The B60 is integrated"** — right conclusion, wrong mechanism, and dangerous
as a generalisation. An earlier revision of this paragraph went further and
called the no-ReBAR case *corruption*; that was wrong, and the correction is
the host-visibility/device-locality distinction above.

**THE SECOND HALF IS THE REAL ONE AND IT IS STILL OWED.** "The gate re-run
where Vulkan actually matters" is now reachable and remains the entire value of
`VK-I`. §0's framing is why: on GB10 *"llama.cpp's own CUDA backend will beat
both of us there. Vulkan on an NVIDIA part is nobody's fastest path; it is the
portability path."* Every Vulkan speed number in this spec is therefore either
llvmpipe (a software rasteriser) or the wrong chip. `garlic-clove` is the first
venue where **Vulkan is the only accelerator path on the box**, so a Vulkan win
here is a win a user would actually feel. Concretely still owed, unchanged:
the 27B prefill/decode re-run and its reference-tier count (named in §6.0b as
"the only thing that can turn the structure above into a result"), `VK-C`'s
coopmat tactic selection on a non-NVIDIA part, and re-taking the
`BENCH-VK-LLAMA` verdict the records call the most fragile in the enumeration
(0.23% margin inside a 0.69% spread, #1003). **20.91 GiB of device-local memory
is the binding constraint** and 27B does not fit; the reachability question is
which model arm does, and the box holds no weights at all.

### 6.0b The DECODE GEMV lever, measured to its floor — 2026-08-09

`row/BACKEND-VULKAN-GEMVROWS`. `vt_matmul_vec` is 89-90% of 27B decode GPU time
Expand Down Expand Up @@ -517,9 +587,16 @@ gated at — and IDENTICAL to what the 128-wide module scores on the same inputs
The residual STREAM is held to the bit-exact tier by `memcmp`, which is what
proves `vt_round_through` is the memory round trip rather than an approximation
of it. `test_vulkan_backend` 30/30 (2371 assertions) on GB10 and 30/30 (1828) on
llvmpipe; `test_opt_paged_engine` with `VLLM_CPP_DEVICE=vulkan` still 6/6
llvmpipe; `test_opt_paged_engine` still 6/6
token-exact (96/96) with 0 declines on BOTH arms; `test_backend_cross_device`
11/11.
11/11. **The device is asserted from the test's own printed `BACKEND PROOF` line
(`the engine selected device type 3`), NOT from an env var — `VLLM_CPP_DEVICE`
is read NOWHERE in the tree.** This line used to read
``test_opt_paged_engine` with `VLLM_CPP_DEVICE=vulkan``, which is a placebo
that happened to select Vulkan only because those builds had no CUDA compiler;
see [benchmark-record.md](../benchmark-record.md) §0 of the
`BACKEND-VULKAN-LOADMEM` entry. The corrected invocation is
[`load-direct-upload.md`](load-direct-upload.md)'s.

**AN HONEST LIMIT ON THAT e2e GATE.** `opt-125m` is a LayerNorm model — two
`vt::LayerNorm` calls per layer and ZERO `vt::RmsNorm` — so the standing
Expand Down Expand Up @@ -714,8 +791,10 @@ and it stays sequential inside the workgroup.
NMSE vs the CPU oracle in the same binary: prefill out `1.47e-14`, prefill
carried state `6.43e-15`, bf16 arm `0`; decode (indexed cache) out `1.64e-14` and
cache `3.31e-15`, decode (compact state) out `1.75e-14`. `test_vulkan_backend`
25/25 cases, 1020/1020 assertions. `test_opt_paged_engine` with
`VLLM_CPP_DEVICE=vulkan` still 6/6 token-exact (96/96), 0 declines.
25/25 cases, 1020/1020 assertions. `test_opt_paged_engine` still 6/6
token-exact (96/96), 0 declines, device asserted from the printed
`BACKEND PROOF` line — `VLLM_CPP_DEVICE` is read nowhere and does not select
it.

**NOT measured: any speed number.** Local Vulkan is llvmpipe. The 27B
prefill/decode re-run on GB10, and the reference-tier count that goes with it,
Expand Down
Loading