Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
4384b77
record: file the 27B capture-write crash issue
lu-zero Sep 28, 2026
77a0571
fix(BACKEND-TENSTORRENT): EnsureDevice2D runs one reshape chain in bo…
lu-zero Sep 28, 2026
34939d9
fix(BACKEND-TENSTORRENT-QWEN35): the fresh-slot Memset installs its s…
lu-zero Sep 28, 2026
1bf953b
fix(BACKEND-TENSTORRENT-QWEN35): the batched decode RAC serves in bot…
lu-zero Sep 28, 2026
8c069f8
fix(BACKEND-TENSTORRENT-QWEN35): the batched decode PA serves in both…
lu-zero Sep 28, 2026
f001440
spec(tt-27b-region-capture): region-scoped decode capture — the 50 Mi…
lu-zero Sep 28, 2026
978fc47
seam(tt-27b-region-capture): the per-region trace-staging census and …
lu-zero Sep 28, 2026
f1c7b01
feat(tt-27b-region-capture): region-scoped decode capture on the 27B …
lu-zero Sep 28, 2026
3eb04a1
record(tt-27b-region-capture): the fit wall measured — 64 live region…
lu-zero Sep 28, 2026
a6bd875
record: close the RAC C=1 issue — the fix landed and the gate evidenc…
lu-zero Sep 28, 2026
5753177
test(tt-27b-region-capture): the handoff census reads its own segment…
lu-zero Sep 28, 2026
9abe733
record(tt-27b-region-capture): the 3.15 GB trace demand is per-comman…
lu-zero Sep 28, 2026
eb17682
test(tt-27b-region-capture): the capture-scope upload guard — red on …
lu-zero Sep 28, 2026
75caf22
fix(tt-27b-region-capture): capture-scope H2D uploads refuse by name …
lu-zero Sep 28, 2026
620a0b2
record(tt-27b-region-capture): the c1 leg falsifies the inline-upload…
lu-zero Sep 28, 2026
492f96a
record(tt-27b-region-capture): attribute the ~3.04 MB per captured co…
lu-zero Sep 28, 2026
d9a5e3c
record(tt-27b-region-capture): the logging discriminator attributes t…
lu-zero Sep 29, 2026
09dc55c
test(tt-27b-region-capture): gate the keepquant region record at the …
lu-zero Sep 29, 2026
2e910b8
fix(tt-27b-region-capture): untrack the build2 artifacts an implement…
lu-zero Oct 1, 2026
8162240
fix(tt-27b-region-capture): launch keepquant with common args only, d…
lu-zero Sep 29, 2026
5951f93
refute(tt-27b-region-capture): stock full-grid op records 17 KB/launc…
lu-zero Sep 29, 2026
0a6c7b0
refute(tt-27b-region-capture): the bisect clears every program-shape …
lu-zero Sep 29, 2026
9288d05
test(tt-27b-region-capture): gate the keepquant region at the 64 KiB …
lu-zero Sep 29, 2026
42496c5
fix(tt-27b-region-capture): untrack the build2 artifacts an implement…
lu-zero Oct 1, 2026
1173669
record(tt-27b-region-capture): re-point the ENG-MOE-HOSTFREE anchor t…
lu-zero Oct 1, 2026
064ae0e
docs(tt-27b-region-capture): declare the region-capture arm's two env…
lu-zero Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/engine-matrix.md

Large diffs are not rendered by default.

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
ID: ISSUE-LOCAL-01M3KM4R2KQN5WXTM57W8BD849
Title: RacIdxCache batched lane mishandles a page-table width change
Row: BACKEND-TENSTORRENT-QWEN35
State: OPEN
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-28
Updated: 2026-09-28
Closed: -

## Problem

WarmRacIdx keys RacIdxCache by (num_slots, block_size) but not page-table width. The C=1 lane reallocates on block_table_cols != e.pt_width (retire + realloc, the #1105 discipline); the batched (num_slots>1) lane added in e39f2cf3f has no such guard: its refresh branch indexes batched_pt_host with the CALLER's block_table_cols against a vector sized at allocation width (OOB read) and copy_to_device's a [1, new_cols] host tensor into a [1, old_cols] device tensor — TT_FATAL 'Host tensor has different shape' (tensor_apis.cpp:161). Exposed by the new batched-PA capture doctest (cols=2) running after the batched-RAC doctest (cols=1) in the full suite.

## Resolution

- 2026-09-28 (worktree row/tt-27b-capture-write) FIXED. The batched lane now
mirrors the C=1 lane's `pt_width` discipline: any `block_table_cols !=
e.batched_pt_width` on a live entry retires the per-user page tables into
`batched_retired_pts` (kept alive — never free a buffer a recorded trace
addresses, #1105), reallocates them at the new width, and resets
`batched_pt_host`; `batched_pt_width` records the allocation width.
`batched_update_idxs` ([1] per user) and the sharded inputs are
width-independent and untouched. Evidence: before the fix the new
batched-PA capture doctest (page-table cols=2) threw TT_FATAL
"Host tensor has different shape" (tensor_apis.cpp:161) when it ran after
the batched-RAC doctest (cols=1) in the full suite — both share RacIdxCache
key (num_slots=2, block_size=32); /tmp/suite-pa.log. After the fix the full
96-case suite is 95/96 with the only failure the pre-existing owed RAC
residual (/tmp/suite-pa3.log: 126/128, user-1 second head — the signature
recorded at e39f2cf3f, unchanged); the PA case reads replay-vs-eagerB
0/2048 mismatched elems in the same full-suite run.
-
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
ID: ISSUE-LOCAL-01M3M0K390EM40W5R9BR5A2KZ7
Title: 27B c1 decode: RAC C=1 lane routes through unallocated batched tensors — segfault at the first cold decode step
Row: BACKEND-TENSTORRENT-QWEN35
State: CLOSED
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-09-28
Updated: 2026-09-28
Closed: 2026-09-28

## Problem

At row/tt-27b-region-capture HEAD ec4e8a824, the Qwen3.8-27B-Q4_K_M c1 leg (--concurrency 1) segfaults in ttnn::copy inside ReshapeAndCacheKernel during the COLD eager decode step (capturing=0; /tmp/leg-control-c1.log, /tmp/leg-region-c1-diag.log, 2026-09-28, thalia). Control leg without VLLM_CPP_REGION_CAPTURE crashes identically, so this is pre-existing on the base, not the region arm. Root cause: e39f2cf3f rewrote TryReshapeAndCacheDeviceDecode as one per-user batched loop (rac_entry.batched_in[u], batched_update_idxs[u], batched_page_table[u]) but WarmRacIdx allocates those ONLY for num_slots>1 — for C=1 it allocates the shared sharded_in/sharded_in_v/update_idxs/page_table and its warm gate admits C=1 on `allocated` alone, so the loop indexes empty vectors (empty ttnn::Tensor -> null storage -> ttnn::Tensor::memory_config() segfault). The commit's claim 'The C=1 lane is untouched' is false; no c1 leg ran on this branch since e39f2cf3f (the doctrine legs were c2). Fix: restore the proven C=1 sequence verbatim (build_input over the whole shadow into the shared sharded tensors, one paged_fused_update_cache against the shared update_idxs/page_table) beside the batched loop.

## Resolution

2026-09-28: fixed in the same flow that found it (commit restoring the C=1 lane verbatim beside the batched loop). Red /tmp/leg-control-c1.log + /tmp/leg-region-c1-diag.log (cold-step segfault, capturing=0, with and without VLLM_CPP_REGION_CAPTURE); green /tmp/leg-region-c1-fix.log (cold step + capture pass run, leg proceeds to the unrelated fit wall) and the full TT suite green on the RAC lane apart from the recorded OWED batched flake. Full detail in the issue's Resolution section.
307 changes: 307 additions & 0 deletions .agents/specs/tt-27b-region-capture.md

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -44,3 +44,4 @@ generated/
# (`651994a03`) re-added all three from another worktree. The ignore has to be
# tracked to hold. Write a PR body OUTSIDE the repository.
.prbody/
build2/
1 change: 1 addition & 0 deletions docs/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,7 @@ These change how the engine runs and have no CLI flag (or complement one).
| `VT_CPU_SPIN_ROUNDS` | `4096` on aarch64, `256` elsewhere | How many relax rounds a CPU-threadpool waiter spins before yielding its core, in `Threadpool::Barrier` and `Threadpool::PollForWork`. A waiter that never yields costs a full scheduler timeslice per dispatch as soon as the pool is wider than the cores available to it, and a stock run reaches that because the pool defaults to hardware concurrency while the process has other runnable threads. `0` restores the never-yield spin for a same-binary A/B. Neither setting changes any computed value: the yield is a scheduling hint only |
| `VLLM_PREFIX_CACHING_HASH_SEED` | `0` (fixed) | Seed for the prefix-cache block hash, mirroring vLLM's `PYTHONHASHSEED`. `random` makes block hashes non-deterministic across processes, which takes any persisted or shared KV cache to a 0% hit rate. Keep it fixed if you rely on cross-process prefix reuse |
| `VLLM_KV_EVENTS_USE_INT_BLOCK_HASHES` | `1` (on) | Whether published KV-cache events carry block hashes as an int (the low 64 bits of the sha256 digest) rather than the raw 32 bytes, mirroring vLLM's env of the same name and its default. Set `0` to publish the raw bytes. Only affects the KV-cache event payload (`--kv-events-config`); it does not change the internal block hashing or the cache itself |
| `VLLM_CPP_REGION_CAPTURE` | `0` (off) | Tenstorrent backend: capture the 27B decode graph as one region per layer instead of one whole graph (`1` enables). Each region is bounded to 50 MiB of trace staging and an over-budget region declines to eager by name; the whole-graph arm is still attempted first when the fit predicate passes. The default path is unchanged — the arm is insurance for the whole-graph trace-fit wall and for devices where whole-graph capture cannot fit |
| `VLLM_PLUGINS` | unset (load all registered) | Comma-separated allowlist of general plugins to load in `LoadGeneralPlugins()`, mirroring vLLM's `VLLM_PLUGINS`. Unset loads every registered plugin; an empty string loads none; a list loads only the named plugins. A plugin that throws is logged and skipped (the load never aborts the engine). See [.agents/specs/plugin-system.md](../.agents/specs/plugin-system.md) |
| `VT_LMCACHE_HOST` | `127.0.0.1` | Default LMCache server host for the `lm://` connector. The `kv_connector_extra_config.host` key overrides it. See [KV offload](KV-OFFLOAD.md) |
| `VT_LMCACHE_PORT` | `65432` | Default LMCache server port. The `kv_connector_extra_config.port` key overrides it |
Expand Down
84 changes: 84 additions & 0 deletions docs/bench-evidence/tt-capture-upload-guard-20260928.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# tt capture-scope upload guard — the leg that falsifies the inline-upload attribution (2026-09-28)

Worktree `row/tt-27b-region-capture-spec`, fixes `286947603` (red test) +
`39e2ca8ef` (guard + broadcast-operand cache), audit input
`docs/bench-evidence/tt-trace-record-audit-20260928.md`. Legs:
`~/.local/logs/maki/CeRgUcXiFK5bYWGSPS4sy/monitor-1790617827-8e22/stdout.log`
(the full suite), `.../monitor-1790618190-57bb/{stdout,stderr}.log` (the 27B
whole-graph c1 leg), and the in-tree region-handoff test run.

## What landed

1. **The guard.** `UploadRows` and `UploadRowsBf16`
(`src/vt/tenstorrent/tenstorrent_residency.cpp`) refuse any H2D upload with
`tt_capture_active()` set, by name, after the required
`VT_TT_TRACE_DEBUG` print on the route. `AddKernel`'s broadcast operand
(`src/vt/tenstorrent/tenstorrent_ops.cpp`) moves behind a cache keyed by
host pointer, geometry, and an FNV-1a hash of the operand's values: the
eager pass uploads once, the capture pass serves the resident copy, and a
capture-scope miss refuses by name (no silent inline, no value staleness).
2. **The red-first test** ("kTENSTORRENT capture-scope upload refuses and the
warmed capture records the 2 KB floor"): red on `286947603`'s parent — the
unwarmed capture-scope upload fired and died on tt-metal's own
`TT_FATAL fd_mesh_command_queue.cpp:826 !trace_id_.has_value()` (no named
refusal) after printing `[TT-UP] UploadRowsBf16 from_span WRITE during
capture rows=1024 cols=1508`. Green after the fix: named refusal + the
warmed capture records **1,024 B**.

## The money leg and what it actually showed

27B whole-graph c1 (`VT_TT_TRACE_DEBUG=1 VT_TT_KEEPQUANT_INT8DOT=0`,
`--num-prompts 2 --input-len 128 --output-len 32 --concurrency 1`):
**BENCH_EXIT=1. No TPOT table — the leg died at the same capture end.**
The fatal is byte-identical to the pre-row one: `end_trace_capture` asks for
**3,153,969,152 B** against 298,568,896 B free
(`bank_manager.cpp` OOM, `assert.hpp:104`).

But the census around it is decisive:

- **Zero** `[TT-UP]` lines in the whole leg: no upload route (staging,
broadcast-Add, rope cache, ids) attempted an H2D write under capture. The
doctrine's eager pass already covers every site — the guard never had to
fire.
- **Zero** `[TT-KQ]` keep-quant word-shadow refusals: that route was already
refusing/covered.
- The 6 `EnsureHostBytes DURING CAPTURE` readbacks fired as before (known
sync hazard, zero trace bytes, still owed).

So on this pin, **the ~3.15 GB demand persists with zero capture-scope
uploads — the audit's model C (inline H2D payload) is falsified for the
default whole-graph arm.** The discriminating experiment the audit itself
proposed settles it at op scale: the region-handoff test with the MatmulBT
fully warmed closes region 1 at exactly **3,088,384 B** — the same close the
audit attributed to a ~3.086 MB inline upload. That payload is NOT an upload;
it is the recorded per-program command stream of the MatmulBT program class
(a warmed RmsNorm region still closes at 2,048 B, so binaries-by-relay holds
for small programs — the ~3 MB rides with the quant-matmul program's launch
record, mechanism unattributed at dispatch level).

## Verdict and next hypothesis

- The guard and the warmable broadcast operand are correct hardening and stay
(they convert the pre-fix raw TT_FATAL into a named refusal and remove the
per-call broadcast re-upload in both passes), but they do not shrink the
whole-graph record, because there was nothing left to shrink from the
upload side.
- The 27B decode-trace DRAM-fit site closes only on a tt-metal-side
attribution: dump the dispatch command stream
(`tt::LogDispatch` trace level) for one captured `MatmulBTQuantGrouped`
launch and find the ~3 MB of `bypass_data` words. Next candidates: the
program's kernel-binary relay pages being re-relayed per launch for
programs that miss the 1,024 KB prefetch ringbuffer
(`fd_mesh_command_queue.cpp:453` fit decision), or per-launch config-page
writes scaling with the quant-program's CB/RTA footprint.
- Region arm unchanged: `VLLM_CPP_REGION_CAPTURE=1` stays env-gated; the
64 live regions summing to the same ~3.15 GB (previous evidence) is now
doubly explained — the demand is per-program, so segmentation cannot help
either.

## Suite

98 cases, 526,778 assertions: **526,777 passed, 1 failed** — the pre-recorded
batched-RAC residual flake (96/128 K/V, user-1 second head,
`test_tenstorrent_backend.cpp:2261`), the exact owed failure the issue
already records. No new failures from the guard.
90 changes: 90 additions & 0 deletions docs/bench-evidence/tt-keepquant-rta-fix-20260929.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
# the per-core RTA fix: measured, and what it did NOT close (2026-09-29)

Worktree `row/tt-27b-region-capture-spec` at the fix commit. Follows
[tt-launch-record-attribution-20260928.md](tt-launch-record-attribution-20260928.md),
whose verdict named our keepquant program's per-core `SetRuntimeArgs`
(`tenstorrent_keepquant.cpp:2108-2130`) as the ~2.9 MB per captured launch.

## 1. The fix (landed)

`src/vt/tenstorrent/tenstorrent_keepquant.cpp`:

- The kernel (`kernel_main`) reads ALL runtime words from
`SetCommonRuntimeArgs` (14 words: the 3 bank bases + M/K/N/nb/wpb/act_f32/
mtile/qb_pad/enc/tcols + grid_x) and derives the per-core slice in-kernel:
`c = get_relative_logical_y() * grid_x + get_relative_logical_x()`,
`row0 = c * tcols`, `rowc = row0 >= N ? 0 : min(tcols, N - row0)` — the
exact guard the deleted host loop applied, including the fully-idle tail.
Grid is part of the workload key, so `grid_x` is shape-global per program.
- The host per-core `SetRuntimeArgs` loop (12 words × grid_cores per call) is
DELETED; every word moves to the common-args vector, set once on a
workload miss and updated in place on a hit. This call site is the only
caller of the program; nothing else needs per-core args on it.

Correctness: the keepquant capture-x2 byte-identity suite stays green on the
fix (E=1 grouped keep-quant capture, full test suite below); the shape
carries a partial last core, so the in-kernel clamp is device-proven. A new
host doctest pins the in-kernel derivation to the deleted loop's values for
every core across partial/idle tail shapes.

## 2. Measured on the 27B whole-graph capture (the arbiter)

c1 leg (`Qwen3.8-27B-Q4_K_M`, 2x128/32, `VT_TT_TRACE_DEBUG=1`), fresh
build2 against the new pin `6449cf13f7b`:

| | trace demand at end_mesh_capture |
|---|---|
| pre-fix (attribution doc) | 3,153,969,152 B |
| post-fix (this leg, /tmp/money-c1.log) | 2,925,109,248 B |

The fix removed **228,859,904 B** ≈ 1,037 captured launches × ~920 grid
cores × one recorded 256 B RTA page each — exactly the per-core
`SetRuntimeArgs` stream the fix deleted. `BENCH_EXIT=1`: the capture still
fatals `Out of Memory: Not enough space to allocate 2925109248 B` (bank
manager, 8 banks). **The whole-graph trace still does not fit DRAM; c1 does
not serve.** The remaining ~2.93 GB is NOT per-core RTAs.

## 3. The attribution doc's per-launch magnitude was wrong; its direction was right

Two controlled A/Bs on small keepquant captures (same command, red vs green
binary, new-pin libs):

- region-handoff doctest: region 1 = **2,965,504 B on BOTH binaries** —
byte-identical. Dispatch-log counts (TT_METAL_LOGGER_LEVEL=TRACE,
build_logging): 22,079 Unique-RTA lines on both.
- E=1 grouped keep-quant capture: device trace demand
**48,316,416 B on BOTH binaries**, capture-x2 byte-identity PASS on both.

So at these shapes the per-core RTA stream contributed ~0 to the recorded
region — the 2.9 MB per captured command is dominated by the ~250 remaining
"one-shot program command sequence" fetches per launch (full-grid CB/DFB
config pages and per-sequence chunks), which are per-launch, not per-core.
The 27B A/B is the honest measurement: −228.9 MB real, wall standing.

## 4. OPEN NEXT (updated)

The per-core RTA lever is SPENT (landed, correct, ~7.3% of the demand). The
dominant remaining class is per-launch program command-sequence payload:
~1,037 launches × ~2.7 MB, i.e. tt-metal records each launch's full-grid
CB/DFB configuration per sequence. Candidate levers, in traceable order:

1. Count the remaining capture-window classes with the existing
logging-enabled discriminator on a 27B capture attempt (the 254-fetch
census of the attribution doc, rerun post-fix) — name the per-sequence
payload composition before touching anything.
2. tt-metal-side: whether `create_trace_node` can dedupe/re-reference
unchanged full-grid config pages across replays of the same program
(upstream question; pin-local experiment first).
3. Our-side: fewer full-grid CB/config-bearing programs per launch (merge
programs), or capture at coarser launch granularity — our-side grid
shrink only scales linearly and stays OOM (recorded as a bound, not a
fix).

## 5. Test reconciliation

The region-handoff KB-bound gate added red-first for this fix measured
2,965,504 B before AND after (§3), so the KB floor is not reachable by
removing per-core RTAs and the gate was removed rather than left red. The
keepquant capture-x2 byte-identity gates and the derivation-parity doctest
stand. The 27B trace-fit assertion remains the bench leg (the only vehicle
at that scale), still failing at 2,925,109,248 B — the row's fit wall.
Loading
Loading