Conversation
FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit b6cc255)
…th passes at the rank-3 arm
The exact-rows/cols arm whose device shadow holds a rank-3 logical shape
under a flat 2D slot record ran a bare ttnn::reshape on the TILED shadow
during capture — a program the eager pass never warmed (its row-major
reshape is a free view) — so the first decode capture created
ReshapeViewTiledProgramFactory's program mid-trace and died on its
unconditional to_device write (ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ).
Both passes now run the warmed to_layout chain. A new doctest reproduces
the exact 'Writes are not supported during trace capture' fatal on the
old branch and passes with the fix; the TT suite is 93/93.
The 27B serve leg still fatals at a second site — a capture-time
same-numel reshape {1,10240}->{2,5120} the eager pass never ran — which
the issue now records; it stays OPEN for that site.
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit ab7cdb3)
…hadow in both passes (issue ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ, site 2)
The 27B serve leg's second capture-write crash came from
MemsetDeviceIfCapture's fresh-slot lane being capture-only for the shadow
install: under capture a fresh DBuf::Zero installed a {1,10240} bf16 TILE
shadow, while the eager pass only primed the zero and kept the host
fallback, so the slot ended each pass in a different state and the
capture-step consumer (kRmsNorm's residual EnsureDevice2D at {2,5120})
hit the same-numel arm with a reshape spec the eager pass never ran.
ReshapeViewTiledProgramFactory created its program mid-trace and its
to_device write fatals at fd_mesh_command_queue.cpp:826. The fix runs
one install in BOTH passes so the eager consumer warms the reshape and
capture replays it as a program-cache hit; eager installs are bounded
to scratch-scale memsets (bytes <= 64 KiB) because an unbounded install
retained the multi-MB weights-load slots and OOMed DRAM. The new doctest
fails red on the old code with the exact bench fatal; the kRopeNeox
(small) case no longer leaks VT_TT_HOST_FREE_DECODE=0 into the rest of
the suite. The 27B leg now passes the site-2 reshape; it next surfaces
two previously masked blockers (multi-slot RAC device path, decode-trace
DRAM fit) recorded in the issue.
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit 9658227)
…h passes — the num_slots>1 decline dies
Blocker A of the 27B capture-write issue. TryReshapeAndCacheDeviceDecode
declined num_slots > 1 ("decode T=1 only for now"), so a c2 serve leg fell
back to the host path inside the capture and EnsureHost(k) read the rope
K/V shadows back mid-trace (TT_FATAL "Reads are not supported during trace
capture", fd_mesh_command_queue.cpp:873). The decline is byte-identical
since 79ff8f3 and was benign at the W3-era c2 bench because rope K/V were
host-resident at the fallback; the capture-warmup redesign's device-
residency waves (28c154d and successors) made them device_current under
capture, turning the perf shortcut into a capture fatal.
The batched lane now runs the proven single-user sequence once per user:
slice the user's [nkv, d] rows out of a fresh native copy of the rope
shadow (rank-3 [C, nkv, d] row slices, or the 27B's rank-2 token-row
[T, nkv*d] via d-aligned per-head column slices), 1.0-multiply, ttnn::copy
into that user's own single-shard persistent input, then one
paged_fused_update_cache per user against per-user [1] update_idx and
[1, cols] page-table tensors WarmRacIdx allocates and refreshes outside
capture. The C=1 lane is untouched. Per-user eager passes sync the queue
so later allocations cannot recycle a prior user's in-flight temporaries.
Red: new doctest reproduces the exact leg fatal with the decline restored
(/tmp/red-multislot.log); green: capture + replay complete and both users'
KV verified token-exact in the paged-KV device shadow via the new
ReadPagedKvShadowForTest hook (/tmp/green-multislot.log). The 27B c2 leg no
longer dies at RAC — it now reaches a FURTHER new site (PagedAttention
multi-slot decline -> host readback during a late re-capture), recorded in
the issue as owed, together with the blocker-B trace-budget analysis
(region-scoped capture recommended; the 3.15 GB whole-graph demand does not
fit). Suite: 94/95 — the new case passes standalone and in subsets; under
the full-suite program-cache history its user-1 second head reads
uninitialized bytes after the eager pass despite verified-correct device
inputs; that residual is recorded in the issue as owed, not hidden.
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
(cherry picked from commit e39f2cf)
… passes — the B>1 capture decline dies Third site of the 27B capture-write class. After e39f2cf the c2 leg served ~16 minutes and then died when a late boundary re-capture reached TryPagedAttentionDeviceDecode's `if (tt_capture_active() && Bu > 1) throw` guard: the B>1 Q 4D materialization declined, the host Q arm refused loudly (VT_CHECK, tenstorrent_paged.cpp:1280), the device path returned false, and PagedAttentionKernel's host oracle EnsureHost-ed the paged-KV cache mid-trace — TT_FATAL "Reads are not supported during trace capture" (fd_mesh_command_queue.cpp:873; diagnosed live with VT_TT_TRACE_DEBUG=1, /tmp/leg-27b-diag.log). The guard is a stale W3 premise: the B>1 arm already runs the IDENTICAL multiply(reshape(...)) chain in both passes, so the eager step warms the free reshape's program for the exact input/output spec and the captured call is a program-cache hit (W4 doctrine, tenstorrent_internal.h CaptureSafeReshape); an unwarmed spec still fatals loudly at the miss. The decline is deleted — one hunk, no capture-active branch. Red: new doctest `kTENSTORRENT batched decode PagedAttention is capture-safe (num_reqs=2)` (mirror of the RAC case) fails on HEAD for the right reason (/tmp/red-pa.log, /tmp/red-pa2.log): it stages generation-B K/V through RAC into the DEVICE paged-KV shadow before the capture, so only a device-served PA can reproduce the reference, and the capture declines (2048/2048 mismatched). Green: capture serves (q_from_device OK cap=1), replay-vs-eagerB 0/2048 (/tmp/green-pa.log). En route the case exposed a real bug in the batched RacIdxCache lane — no page-table width-change guard, OOB host-index read and a shape- mismatch copy_to_device TT_FATAL when another case used a different page-table width under the same (num_slots, block_size) key — fixed by mirroring the C=1 lane's pt_width retire+realloc discipline (ISSUE-LOCAL-01M3KM4R2KQN5WXTM57W8BD849). Suite 95/96 (/tmp/suite-pa3.log); the only failure is the pre-existing owed RAC residual (126/128, user-1 second head), which this does not address. Device gate (/tmp/leg-27b-c2c.log): the PA fatal is gone — the leg served ~12+ minutes through six successful boundary re-captures — and now ends at the already-recorded Blocker B structural limit: end_trace_capture OOMs asking 3,128,655,872 B of trace staging against 278,858,624 B free (bank_manager.cpp:495). That stays with the fresh trace-budget row. No TPOT table; the leg died at a re-capture before the 32-token horizon. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 809067c)
…B per-region budget and the in-place boundary doctrine Whole-graph capture cannot fit the 27B (3.15 GB trace demand vs 2.2 GB best-case supply); one layer per region lands the 9B GDN precedent's 50 MiB budget exactly. One pull request per developer decision 2026-09-28. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit ec4e8a8)
…the whole-graph fit predicate Region-scoped decode capture needs to know what ONE region contributes to the trace buffer, not what 1,037 whole-graph commands accrue. GraphCaptureScope now records a per-segment byte delta from an injectable probe (the Tenstorrent registrar installs LastTraceBytesForTest; a backend that registers none records nothing), BreakableGraph exposes it as region_bytes(), and the pure WholeGraphTraceFits predicate encodes the decline polarity the spec's `## Design` derives from the census. Two host-side cases cover the predicate's arms and the census fill on the recording backend. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit dc99071)
…dense driver, one region per layer The 27B decode graph's whole-graph capture asks end_trace_capture for one ~3.15 GB staging buffer against ~298 MB free and serves nothing; this splits the same command stream one layer per region. The dense decode driver opens its capture kPiecewise under VLLM_CPP_REGION_CAPTURE=1 (the automatic fit-predicate wiring is the next wave), DenseForwardLayers emits a bare GraphBreak after each layer — a pure capture split with no eager call, riding the in-place persistent shadow discipline for the cross-region state — and the driver asserts each region's census against the 50 MiB budget, declining the capture BY NAME (sticky per size, no re-capture storm) when a region is over. Default shape is byte-identical: every break is inert in the kFull arm. The red-first device case captures TWO regions on the real trace backend where region 2 reads the buffer region 1 wrote and proves each replay byte-identical to eager; on the pre-row tree it cannot compile (no census API) and with the env set the whole-graph arm dies at the trace OOM. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit f376b51)
…s sum to the whole graph's staging; the RAC C=1 lane restored Wave-1 evidence (docs/bench-evidence/tt-region-capture-20260928.md): the region arm dies at the 8th segment close on tt-metal mesh_trace.cpp:125 (trace buffer 4,226,469,888 B vs allocation high-water 4,229,506,816 B) — every live region owns its trace staging until release and all 64 must replay each step, so per-layer segmentation does not shrink the ~3.15 GB fit demand. Stop condition recorded; the arm stays masked (env-gated, named decline) and the numbers escalate beside tt-metal#57970. Same flow: ISSUE-LOCAL-01M3M0K390EM40W5R9BR5A2KZ7 — e39f2cf's batched RAC rewrite routed C=1 through batched tensors WarmRacIdx never allocates, segfaulting the first cold decode step of ANY c1 leg (red /tmp/leg-control-c1.log); the proven C=1 sequence is restored verbatim beside the batched loop (green /tmp/leg-region-c1-fix.log). VT_REGION_CENSUS now prints per segment, because a capture that dies mid-scope must still leave its per-region record. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 185a670)
…e is dated The close records the red and green legs and points the suite residual at the standing OWED batched flake, so the issue does not linger open over work its tree already falsifies. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 4485969)
…'s delta, not the process-wide staging level The full TT suite's ~500 prior cases leave the trace-staging byte level draining asynchronously, so region 0's delta in the handoff case can carry a stale subtraction and fail under in-suite history while green standalone (both runs' byte-exactness held). The case now asserts region 1's self-bounded delta against the 50 MiB budget and keeps the byte-identity claims untouched. Suite: 526,771/526,773 — the only failure is the recorded OWED RAC doctest flake (126/128 K elems, user-1 second head), not silently absorbed. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 180beff)
…d inline payload, not fixed per-trace overhead Fits three arithmetic models against the raw evidence (whole-graph 3,153,969,152 B / 1,037 commands, the 2-region census's 2,048 B RmsNorm close and 3,088,384 B MatmulBT close, and the 64-region leg's 4.23 GB high-water). Uniform-per-command is refuted by region 0 (1,487x below the mean); fixed-per-trace is refuted at measured magnitudes (needs 49.3 MB/region vs the 3.09 MB largest measured close). The supported model is ~2 KB dispatch headers plus one inline H2D payload per command where a capture-scope upload fires — 1,037 x 3,088,384 B matches the wall within 1.5%. tt-metal records binaries by reference (relay_paged/relay_ringbuffer, dispatch_settings.cpp:72 ringbuffer), so the ~3 MB term is our lazy EnsureDevice2D/broadcast-Add from_vector staging inside capture; hoisting the upload projects ~2.1 MB total trace and dominates the fewer-traces lever (~395 MB for the biggest 8 layers). FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 6473ae7)
…the pre-fix tree The trace-record audit (docs/bench-evidence/tt-trace-record-audit-20260928.md, cherry-picked from row/tt-trace-record-audit @ 9e28fafa4) attributed the 27B whole-graph 3,153,969,152 B trace demand to inline H2D payloads recorded during capture, not tt-metal record overhead. This case pins the fix contract: an unwarmed EnsureDevice2D upload inside a capture scope must be REFUSED by name, and the warmed capture must record the header floor (~2 KB), not the 3,086,336 B payload the audit measured in a region-1 close. Red evidence (this commit's build, device leg): the unwarmed capture-scope upload fires and dies on tt-metal's own TT_FATAL (fd_mesh_command_queue.cpp:826, "!trace_id_.has_value()") — no named refusal — after printing "[TT-UP] UploadRowsBf16 from_span WRITE during capture rows=1024 cols=1508". The warmed capture already records 1,024 B, confirming the audit's header-floor prediction. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 2869476)
…and the broadcast Add operand warms into a cache The trace-record audit attributed the 27B whole-graph 3,153,969,152 B trace demand to inline H2D payloads recorded during capture. Three routes could fire one: EnsureDevice2D's lazy staging (via UploadRows/UploadRowsBf16), the broadcast AddKernel's unconditional replicated-tensor from_vector, and small per-step ids writes. The capture doctrine (ab7cdb3..809067c) says the eager pass performs every upload and the capture pass finds everything resident. UploadRows and UploadRowsBf16 now REFUSE any upload with tt_capture_active() set, naming the site and the warm-before-TraceBeginCapture contract (the VT_TT_TRACE_DEBUG print on the route stays, so a leg proves zero uploads under capture). AddKernel's broadcast operand moves behind a cache keyed by host pointer, geometry, and an FNV-1a hash of the d operand values: the eager pass uploads once, the capture pass serves the resident copy, a changed operand re-uploads in the next eager pass, and a capture-scope miss refuses by name — no silent inline payload and no value-staleness. The focused case (previous commit) goes green: the unwarmed capture-scope upload refuses with the named error, and the warmed capture records 1,024 B — the header floor, against the 3,086,336 B payload the audit measured. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 39e2ca8)
… attribution — the ~3 MB per command is the quant-matmul program's own launch stream The 27B whole-graph c1 leg (VT_TT_TRACE_DEBUG=1) served zero capture-scope uploads — no [TT-UP] route print, no keep-quant refusal — and end_trace_capture still demanded byte-identical 3,153,969,152 B (BENCH_EXIT=1, no TPOT table). The audit's own discriminating experiment settles the attribution at op scale: the region-handoff MatmulBT, fully warmed, closes its region at exactly 3,088,384 B — the close the audit had attributed to a ~3.086 MB inline upload. The payload is the quant-matmul program class's recorded launch stream on this tt-metal pin, not a capture-scope upload; the warmed RmsNorm region still closes at 2,048 B, so binaries-by-relay holds for small programs. The wave's code stays: the guard converts the pre-fix raw TT_FATAL (fd_mesh_command_queue.cpp:826) into a named refusal and the broadcast-Add operand warms into a cache instead of re-uploading in both passes (red 2869476, green 39e2ca8; suite 98 cases / 526,777 of 526,778 assertions green, the only failure the pre-recorded RAC residual flake). The spec's ## Now and the row issue record the falsification and the sharpened next step: dump the tt::LogDispatch command stream for one captured MatmulBTQuantGrouped launch and attribute the ~3 MB of bypass_data. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 5f7ac1f)
…mmand to tt-metal's per-core launch record, not binary relay The binary-relay candidate is falsified at source: the 1,024 KB prefetch-ringbuffer fit only chooses relay_paged vs relay_ringbuffer and both record the kernel binary by reference to the resident DRAM kernels buffer, so binary bytes never enter the trace, and load_binaries refuses first-time loads mid-capture by name; the compiled keep-quant kernel measures 51,648 B of text against that 1,024 KB threshold. Device legs under gdb breakpins on the region-handoff doctest show the 3,088,384 B region record is 284 issue-queue chunks of the recorded command stream for the one full-grid MatmulBTQuant program with zero in-capture buffer-data writes, so the scaling class is the per-core config/RTA dispatch footprint — a tt-metal-side fix, with the one remaining per-chunk discriminator named for a logging-enabled pin build. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 4ca44ee)
…he 2.9 MB region record to our per-core RTA stream The tt-metal pin advanced to upstream-live 98134127a7b (pin head 6449cf13f7b, branch vllm-cpp-pin/20260923-adv; same local series rebased, three rebase resolutions recorded there). With TT_METAL_ENABLE_LOGGING=ON in the pin's build_logging dir the focused leg reads region 0 = 2,048 B, region 1 = 2,965,504 B, and the capture window is exactly 254 one-shot command-sequence fetches (2,961,024 B) holding 11,040 per-core Unique RTA unicast writes — our keepquant SetRuntimeArgs stream (tenstorrent_keepquant.cpp:2108-2130), 12 words per core of which only r0/rc vary and both are kernel-derivable. Verdict flips to ours: compute r0/rc in-kernel and launch with SetCommonRuntimeArgs only; expected record drops to the RmsNorm-class floor (~200x), closing the 27B trace-fit site. Suite on the new pin: 97/98 cases, 526,777/526,778 assertions (the token-exact K flake), plus a post-summary teardown segfault in ttnn::Tensor::deallocate_impl -> GraphTracker::is_enabled to chase upstream. The chunked GDN mirror in tenstorrent_internal.h gains the use_mcast parameter the advanced pin inserted at position 12, which the row build needs to link against the new pin. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 0b6f434)
…KB floor The region-handoff case gains the 27B trace-fit bound: region 1's MatmulBT is the keepquant program over the full worker grid, and on the per-core SetRuntimeArgs tree the captured launch records 2,965,504 B — the bound (65536 B) fails on HEAD with that measured number, so the case is RED before the fix lands. A host-side case pins the in-kernel row0/rowc derivation to the deleted per-core loop's values for every core, including partial and idle tail cores. The device arm of that proof is the region-handoff byte-exact replay, whose shape carries a partial last core. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit a166111)
…erive r0/rc in-kernel The keepquant int8-dot program's per-core SetRuntimeArgs loop (12 words per core per call) is deleted: the kernel now reads all runtime words from SetCommonRuntimeArgs (14 shape-global words including grid_x, which the workload key already pins per program) and derives its column slice from the core coordinate as c = y*grid_x + x, row0 = c*tcols, rowc = the same clamp the deleted loop applied, with fully-idle tail cores returning at the existing rowc == 0 guard. Measured on the 27B whole-graph c1 capture (fresh build2, pin 6449cf13f7b): trace demand drops 3,153,969,152 -> 2,925,109,248 B (-228,859,904 B = 1,037 launches x ~920 cores x one recorded 256 B RTA page), but the trace still exceeds DRAM and BENCH_EXIT=1 — the fit wall stands on a different, per-launch record class. Controlled A/Bs on small keepquant captures (region-handoff 2,965,504 B both binaries; grouped capture 48,316,416 B both binaries) falsify the per-launch ~2.9 MB attribution to the RTA stream at those shapes, so the red-first KB-floor gate from a166111 is removed as unreachable by this lever and the falsification is recorded. Keepquant capture-x2 byte-identity stays green; a host doctest pins the in-kernel derivation to the deleted loop's values across partial and idle tail shapes. Evidence and the updated next lever: docs/bench-evidence/tt-keepquant-rta-fix-20260929.md; issue ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ carries the dated entry. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit ae27ec7)
…h, config-page scaling is program-specific Standalone reproducer (no vllm.cpp) against pin 6449cf13f7b on the P150: ttnn::multiply captured in a trace at a 1-core extent and at the full 110-core height-sharded grid records byte-identical 17,408 B, and an 8-launch trace is exactly 8x, so per-launch trace cost is grid-independent for stock ops. The generic per-launch config-pages x grid hypothesis for the 27B residual (~2.82 MB x 1,037 launches post RTA fix) is refuted as an upstream ask; no tt-metal issue filed. The residual must come from the keepquant program's own config structure (kernel/CB page count or unpacked pages); the evidence note lists the next bisect leg. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 600bfd6)
…knob, the record is grouped-decode program count The synthetic harness builds the keepquant program's exact shape raw (full-grid CoreRange data-movement kernel, common runtime args, CBs) and sweeps CB count, CB page size and kernel binary size: every variant records 1,024 B per traced launch, so the packed relay already collapses identical per-core config pages and even 256 KiB binaries ride free. The 2,965,504 B region record on HEAD is instead ~680 small programs per captured region: the W4a grouped arm (the default dispatch) runs the Q6_K eltwise word decode ~85 programs per chunk times ceil(N/8) chunks. The next lever is collapsing that decode to one custom-kernel program — the bit-exact int8-dot kernel is the existence proof at the floor; the 64 KiB doctest gate landed in the next commit is its arbiter. Evidence note and the row issue carry the numbers. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit 5945879)
…floor — RED at the measured 2,965,504 B The region-handoff census tightens region 1 from the 50 MiB staging budget to the config-page floor the repro proved (a full-grid program records 1,024-17,408 B per launch). On HEAD 600bfd6 the gate reads 2,965,504 B (/tmp/region-red.log) and fails, because the default grouped dispatch captures ~680 decode-chain programs per region. The gate is the arbiter for the owed decode-fusion lever and stays RED until that lands. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki] (cherry picked from commit f92eae3)
…ation commit swept in Build outputs rode along in the RTA-fix commit; they are not sources. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
lu-zero
force-pushed
the
row/tt-27b-region-capture
branch
from
September 30, 2026 21:19
f92eae3 to
bac51b1
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The 27B Tenstorrent decode arm has been unrunnable at main since the capture-warmup redesign: the whole-graph trace could not fit DRAM (3,153,969,152 B demanded, ~298 MB free), and the first crash on the path was a capture-semantics fatal. This branch carries the full investigation to that wall, the fixes found along the way, and the instrumentation the next row will use to finish the job.
What landed:
Known owed: the full-suite order-sensitive RAC doctest residual (126==128, user-1 second head; recorded, not masked); the 27B serve arm stays blocked on the fusion row; the INT8DOT A/B re-measure, golden adjudication, and W1 GEMV legs queue behind it. The device-gated doctests do not run in CI; they were gated green on the P150 (526,777/526,778 + the recorded flake).
Issues: ISSUE-LOCAL-01M3JXEFQKSZP23PP2HWY9G0VQ (open, carries the full attribution), ISSUE-LOCAL-01M3M0K390EM40W5R9BR5A2KZ (closed), ISSUE-LOCAL-01M3KM4R2KQN5WXTM57W8BD849 (closed).
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]