Skip to content

fix(BACKEND-TENSTORRENT): the 27B capture row — the crash onion dies, the fit wall is attributed, the fusion gate is set - #3362

Open
lu-zero wants to merge 23 commits into
mudler:mainfrom
lu-zero:row/tt-27b-region-capture
Open

lu-zero wants to merge 23 commits into
mudler:mainfrom
lu-zero:row/tt-27b-region-capture

Conversation

@lu-zero

@lu-zero lu-zero commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

The 27B Tenstorrent decode arm has been unrunnable at main since the capture-warmup redesign: the whole-graph trace could not fit DRAM (3,153,969,152 B demanded, ~298 MB free), and the first crash on the path was a capture-semantics fatal. This branch carries the full investigation to that wall, the fixes found along the way, and the instrumentation the next row will use to finish the job.

What landed:

  • Four capture-semantics fixes (the crash onion): EnsureDevice2D runs one warmed reshape chain in both passes (rank-3 arm); MemsetDeviceIfCapture installs the fresh-slot shadow in both passes (bounded to scratch-scale memsets); batched decode RAC admitted with a per-user device path; the stale W3 B>1 PagedAttention capture guard deleted. Each with a red-first doctest; the C=1 RAC lane a batching commit had silently swallowed is restored (own issue, closed).
  • Capture-scope H2D uploads now refuse by name (UploadRows/UploadRowsBf16 guard, broadcast Add operand warmed into a device cache) — permanent hardening, verified zero fires in the full leg.
  • Keepquant launches common-runtime-args only; the kernel derives its per-core slice in-kernel, bit-exact including the partial last core (-229 MB of trace demand, measured).
  • The BreakableGraph seam: per-segment trace-staging census (VT_REGION_CENSUS), the WholeGraphTraceFits predicate, one-region-per-layer capture on the 27B driver (env-gated, named per-size decline). The region arm is insurance; the census is the durable instrument.
  • Evidence and instruments: the standalone trace-size harness, the program-shape bisect, and five dated bench-evidence notes that decompose the fit wall completely. The wall is now attributed to program count in the grouped decode chain (~680 small programs per region x the 4-17 KB per-program floor); the decode-fusion row that collapses it is specced as next, with its red gate already in the tree (tests/vt/test_tenstorrent_backend.cpp:11527, intentionally failing on device, named in the issue).
  • tt-metal pin advanced to 6449cf13f7b (live main 98134127a7b + rebased patch series) with the GDN signature adaptation mirrored.

Known owed: the full-suite order-sensitive RAC doctest residual (126==128, user-1 second head; recorded, not masked); the 27B serve arm stays blocked on the fusion row; the INT8DOT A/B re-measure, golden adjudication, and W1 GEMV legs queue behind it. The device-gated doctests do not run in CI; they were gated green on the P150 (526,777/526,778 + the recorded flake).

Issues: ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ (open, carries the full attribution), ISSUE-LOCAL-01M3M0K390EM40W5R9BR5A2KZ (closed), ISSUE-LOCAL-01M3KM4R2KQN5WXTM57W8BD849 (closed).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit b6cc255)
…th passes at the rank-3 arm

The exact-rows/cols arm whose device shadow holds a rank-3 logical shape
under a flat 2D slot record ran a bare ttnn::reshape on the TILED shadow
during capture — a program the eager pass never warmed (its row-major
reshape is a free view) — so the first decode capture created
ReshapeViewTiledProgramFactory's program mid-trace and died on its
unconditional to_device write (ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ).
Both passes now run the warmed to_layout chain. A new doctest reproduces
the exact 'Writes are not supported during trace capture' fatal on the
old branch and passes with the fix; the TT suite is 93/93.

The 27B serve leg still fatals at a second site — a capture-time
same-numel reshape {1,10240}->{2,5120} the eager pass never ran — which
the issue now records; it stays OPEN for that site.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit ab7cdb3)
…hadow in both passes (issue ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ, site 2)

The 27B serve leg's second capture-write crash came from
MemsetDeviceIfCapture's fresh-slot lane being capture-only for the shadow
install: under capture a fresh DBuf::Zero installed a {1,10240} bf16 TILE
shadow, while the eager pass only primed the zero and kept the host
fallback, so the slot ended each pass in a different state and the
capture-step consumer (kRmsNorm's residual EnsureDevice2D at {2,5120})
hit the same-numel arm with a reshape spec the eager pass never ran.
ReshapeViewTiledProgramFactory created its program mid-trace and its
to_device write fatals at fd_mesh_command_queue.cpp:826. The fix runs
one install in BOTH passes so the eager consumer warms the reshape and
capture replays it as a program-cache hit; eager installs are bounded
to scratch-scale memsets (bytes <= 64 KiB) because an unbounded install
retained the multi-MB weights-load slots and OOMed DRAM. The new doctest
fails red on the old code with the exact bench fatal; the kRopeNeox
(small) case no longer leaks VT_TT_HOST_FREE_DECODE=0 into the rest of
the suite. The 27B leg now passes the site-2 reshape; it next surfaces
two previously masked blockers (multi-slot RAC device path, decode-trace
DRAM fit) recorded in the issue.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 9658227)
…h passes — the num_slots>1 decline dies

Blocker A of the 27B capture-write issue. TryReshapeAndCacheDeviceDecode
declined num_slots > 1 ("decode T=1 only for now"), so a c2 serve leg fell
back to the host path inside the capture and EnsureHost(k) read the rope
K/V shadows back mid-trace (TT_FATAL "Reads are not supported during trace
capture", fd_mesh_command_queue.cpp:873). The decline is byte-identical
since 79ff8f3 and was benign at the W3-era c2 bench because rope K/V were
host-resident at the fallback; the capture-warmup redesign's device-
residency waves (28c154d and successors) made them device_current under
capture, turning the perf shortcut into a capture fatal.

The batched lane now runs the proven single-user sequence once per user:
slice the user's [nkv, d] rows out of a fresh native copy of the rope
shadow (rank-3 [C, nkv, d] row slices, or the 27B's rank-2 token-row
[T, nkv*d] via d-aligned per-head column slices), 1.0-multiply, ttnn::copy
into that user's own single-shard persistent input, then one
paged_fused_update_cache per user against per-user [1] update_idx and
[1, cols] page-table tensors WarmRacIdx allocates and refreshes outside
capture. The C=1 lane is untouched. Per-user eager passes sync the queue
so later allocations cannot recycle a prior user's in-flight temporaries.

Red: new doctest reproduces the exact leg fatal with the decline restored
(/tmp/red-multislot.log); green: capture + replay complete and both users'
KV verified token-exact in the paged-KV device shadow via the new
ReadPagedKvShadowForTest hook (/tmp/green-multislot.log). The 27B c2 leg no
longer dies at RAC — it now reaches a FURTHER new site (PagedAttention
multi-slot decline -> host readback during a late re-capture), recorded in
the issue as owed, together with the blocker-B trace-budget analysis
(region-scoped capture recommended; the 3.15 GB whole-graph demand does not
fit). Suite: 94/95 — the new case passes standalone and in subsets; under
the full-suite program-cache history its user-1 second head reads
uninitialized bytes after the eager pass despite verified-correct device
inputs; that residual is recorded in the issue as owed, not hidden.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit e39f2cf)
… passes — the B>1 capture decline dies

Third site of the 27B capture-write class. After e39f2cf the c2 leg
served ~16 minutes and then died when a late boundary re-capture reached
TryPagedAttentionDeviceDecode's `if (tt_capture_active() && Bu > 1)
throw` guard: the B>1 Q 4D materialization declined, the host Q arm
refused loudly (VT_CHECK, tenstorrent_paged.cpp:1280), the device path
returned false, and PagedAttentionKernel's host oracle EnsureHost-ed the
paged-KV cache mid-trace — TT_FATAL "Reads are not supported during
trace capture" (fd_mesh_command_queue.cpp:873; diagnosed live with
VT_TT_TRACE_DEBUG=1, /tmp/leg-27b-diag.log).

The guard is a stale W3 premise: the B>1 arm already runs the IDENTICAL
multiply(reshape(...)) chain in both passes, so the eager step warms the
free reshape's program for the exact input/output spec and the captured
call is a program-cache hit (W4 doctrine, tenstorrent_internal.h
CaptureSafeReshape); an unwarmed spec still fatals loudly at the miss.
The decline is deleted — one hunk, no capture-active branch.

Red: new doctest `kTENSTORRENT batched decode PagedAttention is
capture-safe (num_reqs=2)` (mirror of the RAC case) fails on HEAD for
the right reason (/tmp/red-pa.log, /tmp/red-pa2.log): it stages
generation-B K/V through RAC into the DEVICE paged-KV shadow before the
capture, so only a device-served PA can reproduce the reference, and
the capture declines (2048/2048 mismatched). Green: capture serves
(q_from_device OK cap=1), replay-vs-eagerB 0/2048 (/tmp/green-pa.log).
En route the case exposed a real bug in the batched RacIdxCache lane —
no page-table width-change guard, OOB host-index read and a shape-
mismatch copy_to_device TT_FATAL when another case used a different
page-table width under the same (num_slots, block_size) key — fixed by
mirroring the C=1 lane's pt_width retire+realloc discipline
(ISSUE-LOCAL-01M3KM4R2KQN5WXTM57W8BD849). Suite 95/96
(/tmp/suite-pa3.log); the only failure is the pre-existing owed RAC
residual (126/128, user-1 second head), which this does not address.

Device gate (/tmp/leg-27b-c2c.log): the PA fatal is gone — the leg
served ~12+ minutes through six successful boundary re-captures — and
now ends at the already-recorded Blocker B structural limit:
end_trace_capture OOMs asking 3,128,655,872 B of trace staging against
278,858,624 B free (bank_manager.cpp:495). That stays with the fresh
trace-budget row. No TPOT table; the leg died at a re-capture before
the 32-token horizon.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 809067c)
…B per-region budget and the in-place boundary doctrine

Whole-graph capture cannot fit the 27B (3.15 GB trace demand vs 2.2 GB best-case supply); one layer per region lands the 9B GDN precedent's 50 MiB budget exactly. One pull request per developer decision 2026-09-28.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit ec4e8a8)
…the whole-graph fit predicate

Region-scoped decode capture needs to know what ONE region contributes to the
trace buffer, not what 1,037 whole-graph commands accrue. GraphCaptureScope now
records a per-segment byte delta from an injectable probe (the Tenstorrent
registrar installs LastTraceBytesForTest; a backend that registers none records
nothing), BreakableGraph exposes it as region_bytes(), and the pure
WholeGraphTraceFits predicate encodes the decline polarity the spec's
`## Design` derives from the census. Two host-side cases cover the predicate's
arms and the census fill on the recording backend.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit dc99071)
…dense driver, one region per layer

The 27B decode graph's whole-graph capture asks end_trace_capture for one
~3.15 GB staging buffer against ~298 MB free and serves nothing; this splits the
same command stream one layer per region. The dense decode driver opens its
capture kPiecewise under VLLM_CPP_REGION_CAPTURE=1 (the automatic fit-predicate
wiring is the next wave), DenseForwardLayers emits a bare GraphBreak after each
layer — a pure capture split with no eager call, riding the in-place persistent
shadow discipline for the cross-region state — and the driver asserts each
region's census against the 50 MiB budget, declining the capture BY NAME (sticky
per size, no re-capture storm) when a region is over. Default shape is
byte-identical: every break is inert in the kFull arm. The red-first device case
captures TWO regions on the real trace backend where region 2 reads the buffer
region 1 wrote and proves each replay byte-identical to eager; on the pre-row
tree it cannot compile (no census API) and with the env set the whole-graph arm
dies at the trace OOM.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit f376b51)
…s sum to the whole graph's staging; the RAC C=1 lane restored

Wave-1 evidence (docs/bench-evidence/tt-region-capture-20260928.md): the
region arm dies at the 8th segment close on tt-metal mesh_trace.cpp:125
(trace buffer 4,226,469,888 B vs allocation high-water 4,229,506,816 B) —
every live region owns its trace staging until release and all 64 must
replay each step, so per-layer segmentation does not shrink the ~3.15 GB
fit demand. Stop condition recorded; the arm stays masked (env-gated,
named decline) and the numbers escalate beside tt-metal#57970. Same flow:
ISSUE-LOCAL-01M3M0K390EM40W5R9BR5A2KZ7 — e39f2cf's batched RAC rewrite
routed C=1 through batched tensors WarmRacIdx never allocates, segfaulting
the first cold decode step of ANY c1 leg (red /tmp/leg-control-c1.log);
the proven C=1 sequence is restored verbatim beside the batched loop
(green /tmp/leg-region-c1-fix.log). VT_REGION_CENSUS now prints per
segment, because a capture that dies mid-scope must still leave its
per-region record.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 185a670)
…e is dated

The close records the red and green legs and points the suite residual at the
standing OWED batched flake, so the issue does not linger open over work its
tree already falsifies.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 4485969)
…'s delta, not the process-wide staging level

The full TT suite's ~500 prior cases leave the trace-staging byte level
draining asynchronously, so region 0's delta in the handoff case can carry a
stale subtraction and fail under in-suite history while green standalone (both
runs' byte-exactness held). The case now asserts region 1's self-bounded delta
against the 50 MiB budget and keeps the byte-identity claims untouched. Suite:
526,771/526,773 — the only failure is the recorded OWED RAC doctest flake
(126/128 K elems, user-1 second head), not silently absorbed.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 180beff)
…d inline payload, not fixed per-trace overhead

Fits three arithmetic models against the raw evidence (whole-graph
3,153,969,152 B / 1,037 commands, the 2-region census's 2,048 B RmsNorm
close and 3,088,384 B MatmulBT close, and the 64-region leg's 4.23 GB
high-water). Uniform-per-command is refuted by region 0 (1,487x below the
mean); fixed-per-trace is refuted at measured magnitudes (needs
49.3 MB/region vs the 3.09 MB largest measured close). The supported model
is ~2 KB dispatch headers plus one inline H2D payload per command where a
capture-scope upload fires — 1,037 x 3,088,384 B matches the wall within
1.5%. tt-metal records binaries by reference (relay_paged/relay_ringbuffer,
dispatch_settings.cpp:72 ringbuffer), so the ~3 MB term is our lazy
EnsureDevice2D/broadcast-Add from_vector staging inside capture; hoisting
the upload projects ~2.1 MB total trace and dominates the fewer-traces
lever (~395 MB for the biggest 8 layers).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 6473ae7)
…the pre-fix tree

The trace-record audit (docs/bench-evidence/tt-trace-record-audit-20260928.md,
cherry-picked from row/tt-trace-record-audit @ 9e28fafa4) attributed the 27B
whole-graph 3,153,969,152 B trace demand to inline H2D payloads recorded during
capture, not tt-metal record overhead. This case pins the fix contract: an
unwarmed EnsureDevice2D upload inside a capture scope must be REFUSED by name,
and the warmed capture must record the header floor (~2 KB), not the
3,086,336 B payload the audit measured in a region-1 close.

Red evidence (this commit's build, device leg): the unwarmed capture-scope
upload fires and dies on tt-metal's own TT_FATAL
(fd_mesh_command_queue.cpp:826, "!trace_id_.has_value()") — no named refusal —
after printing "[TT-UP] UploadRowsBf16 from_span WRITE during capture
rows=1024 cols=1508". The warmed capture already records 1,024 B, confirming
the audit's header-floor prediction.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 2869476)
…and the broadcast Add operand warms into a cache

The trace-record audit attributed the 27B whole-graph 3,153,969,152 B trace
demand to inline H2D payloads recorded during capture. Three routes could fire
one: EnsureDevice2D's lazy staging (via UploadRows/UploadRowsBf16), the
broadcast AddKernel's unconditional replicated-tensor from_vector, and small
per-step ids writes. The capture doctrine (ab7cdb3..809067c) says the eager
pass performs every upload and the capture pass finds everything resident.

UploadRows and UploadRowsBf16 now REFUSE any upload with tt_capture_active()
set, naming the site and the warm-before-TraceBeginCapture contract (the
VT_TT_TRACE_DEBUG print on the route stays, so a leg proves zero uploads under
capture). AddKernel's broadcast operand moves behind a cache keyed by host
pointer, geometry, and an FNV-1a hash of the d operand values: the eager pass
uploads once, the capture pass serves the resident copy, a changed operand
re-uploads in the next eager pass, and a capture-scope miss refuses by name —
no silent inline payload and no value-staleness.

The focused case (previous commit) goes green: the unwarmed capture-scope
upload refuses with the named error, and the warmed capture records 1,024 B —
the header floor, against the 3,086,336 B payload the audit measured.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 39e2ca8)
… attribution — the ~3 MB per command is the quant-matmul program's own launch stream

The 27B whole-graph c1 leg (VT_TT_TRACE_DEBUG=1) served zero capture-scope
uploads — no [TT-UP] route print, no keep-quant refusal — and end_trace_capture
still demanded byte-identical 3,153,969,152 B (BENCH_EXIT=1, no TPOT table).
The audit's own discriminating experiment settles the attribution at op scale:
the region-handoff MatmulBT, fully warmed, closes its region at exactly
3,088,384 B — the close the audit had attributed to a ~3.086 MB inline upload.
The payload is the quant-matmul program class's recorded launch stream on this
tt-metal pin, not a capture-scope upload; the warmed RmsNorm region still
closes at 2,048 B, so binaries-by-relay holds for small programs.

The wave's code stays: the guard converts the pre-fix raw TT_FATAL
(fd_mesh_command_queue.cpp:826) into a named refusal and the broadcast-Add
operand warms into a cache instead of re-uploading in both passes (red
2869476, green 39e2ca8; suite 98 cases / 526,777 of 526,778 assertions
green, the only failure the pre-recorded RAC residual flake). The spec's
## Now and the row issue record the falsification and the sharpened next
step: dump the tt::LogDispatch command stream for one captured
MatmulBTQuantGrouped launch and attribute the ~3 MB of bypass_data.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 5f7ac1f)
…mmand to tt-metal's per-core launch record, not binary relay

The binary-relay candidate is falsified at source: the 1,024 KB
prefetch-ringbuffer fit only chooses relay_paged vs relay_ringbuffer and both
record the kernel binary by reference to the resident DRAM kernels buffer, so
binary bytes never enter the trace, and load_binaries refuses first-time loads
mid-capture by name; the compiled keep-quant kernel measures 51,648 B of text
against that 1,024 KB threshold. Device legs under gdb breakpins on the
region-handoff doctest show the 3,088,384 B region record is 284 issue-queue
chunks of the recorded command stream for the one full-grid MatmulBTQuant
program with zero in-capture buffer-data writes, so the scaling class is the
per-core config/RTA dispatch footprint — a tt-metal-side fix, with the one
remaining per-chunk discriminator named for a logging-enabled pin build.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 4ca44ee)
…he 2.9 MB region record to our per-core RTA stream

The tt-metal pin advanced to upstream-live 98134127a7b (pin head
6449cf13f7b, branch vllm-cpp-pin/20260923-adv; same local series rebased,
three rebase resolutions recorded there). With TT_METAL_ENABLE_LOGGING=ON in
the pin's build_logging dir the focused leg reads region 0 = 2,048 B,
region 1 = 2,965,504 B, and the capture window is exactly 254 one-shot
command-sequence fetches (2,961,024 B) holding 11,040 per-core Unique RTA
unicast writes — our keepquant SetRuntimeArgs stream
(tenstorrent_keepquant.cpp:2108-2130), 12 words per core of which only
r0/rc vary and both are kernel-derivable. Verdict flips to ours: compute
r0/rc in-kernel and launch with SetCommonRuntimeArgs only; expected record
drops to the RmsNorm-class floor (~200x), closing the 27B trace-fit site.
Suite on the new pin: 97/98 cases, 526,777/526,778 assertions (the token-exact
K flake), plus a post-summary teardown segfault in
ttnn::Tensor::deallocate_impl -> GraphTracker::is_enabled to chase upstream.
The chunked GDN mirror in tenstorrent_internal.h gains the use_mcast
parameter the advanced pin inserted at position 12, which the row build
needs to link against the new pin.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 0b6f434)
…KB floor

The region-handoff case gains the 27B trace-fit bound: region 1's MatmulBT
is the keepquant program over the full worker grid, and on the per-core
SetRuntimeArgs tree the captured launch records 2,965,504 B — the bound
(65536 B) fails on HEAD with that measured number, so the case is RED
before the fix lands. A host-side case pins the in-kernel row0/rowc
derivation to the deleted per-core loop's values for every core, including
partial and idle tail cores. The device arm of that proof is the
region-handoff byte-exact replay, whose shape carries a partial last core.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit a166111)
…erive r0/rc in-kernel

The keepquant int8-dot program's per-core SetRuntimeArgs loop (12 words per
core per call) is deleted: the kernel now reads all runtime words from
SetCommonRuntimeArgs (14 shape-global words including grid_x, which the
workload key already pins per program) and derives its column slice from the
core coordinate as c = y*grid_x + x, row0 = c*tcols, rowc = the same clamp
the deleted loop applied, with fully-idle tail cores returning at the
existing rowc == 0 guard. Measured on the 27B whole-graph c1 capture (fresh
build2, pin 6449cf13f7b): trace demand drops 3,153,969,152 -> 2,925,109,248
B (-228,859,904 B = 1,037 launches x ~920 cores x one recorded 256 B RTA
page), but the trace still exceeds DRAM and BENCH_EXIT=1 — the fit wall
stands on a different, per-launch record class. Controlled A/Bs on small
keepquant captures (region-handoff 2,965,504 B both binaries; grouped
capture 48,316,416 B both binaries) falsify the per-launch ~2.9 MB
attribution to the RTA stream at those shapes, so the red-first KB-floor
gate from a166111 is removed as unreachable by this lever and the
falsification is recorded. Keepquant capture-x2 byte-identity stays green;
a host doctest pins the in-kernel derivation to the deleted loop's values
across partial and idle tail shapes. Evidence and the updated next lever:
docs/bench-evidence/tt-keepquant-rta-fix-20260929.md; issue
ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ carries the dated entry.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit ae27ec7)
…h, config-page scaling is program-specific

Standalone reproducer (no vllm.cpp) against pin 6449cf13f7b on the P150:
ttnn::multiply captured in a trace at a 1-core extent and at the full
110-core height-sharded grid records byte-identical 17,408 B, and an
8-launch trace is exactly 8x, so per-launch trace cost is grid-independent
for stock ops. The generic per-launch config-pages x grid hypothesis for
the 27B residual (~2.82 MB x 1,037 launches post RTA fix) is refuted as an
upstream ask; no tt-metal issue filed. The residual must come from the
keepquant program's own config structure (kernel/CB page count or unpacked
pages); the evidence note lists the next bisect leg.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 600bfd6)
…knob, the record is grouped-decode program count

The synthetic harness builds the keepquant program's exact shape raw
(full-grid CoreRange data-movement kernel, common runtime args, CBs) and
sweeps CB count, CB page size and kernel binary size: every variant records
1,024 B per traced launch, so the packed relay already collapses identical
per-core config pages and even 256 KiB binaries ride free. The 2,965,504 B
region record on HEAD is instead ~680 small programs per captured region:
the W4a grouped arm (the default dispatch) runs the Q6_K eltwise word
decode ~85 programs per chunk times ceil(N/8) chunks. The next lever is
collapsing that decode to one custom-kernel program — the bit-exact
int8-dot kernel is the existence proof at the floor; the 64 KiB doctest
gate landed in the next commit is its arbiter. Evidence note and the row
issue carry the numbers.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 5945879)
…floor — RED at the measured 2,965,504 B

The region-handoff census tightens region 1 from the 50 MiB staging budget
to the config-page floor the repro proved (a full-grid program records
1,024-17,408 B per launch). On HEAD 600bfd6 the gate reads 2,965,504 B
(/tmp/region-red.log) and fails, because the default grouped dispatch
captures ~680 decode-chain programs per region. The gate is the arbiter
for the owed decode-fusion lever and stays RED until that lands.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit f92eae3)
…ation commit swept in

Build outputs rode along in the RTA-fix commit; they are not sources.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
@lu-zero
lu-zero force-pushed the row/tt-27b-region-capture branch from f92eae3 to bac51b1 Compare September 30, 2026 21:19

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant