Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion packages/acestep/optimization/LEDGER.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
Stage 2 was explicitly authorized on 2026-08-13. The approved baseline is
frozen and measured optimization is active.

Next available ID: `OPT-0089`.
Next available ID: `OPT-0090`.

| ID | Subsystem | Hypothesis | Evidence | Disposition | Result | Record | Implementation |
| --- | --- | --- | --- | --- | --- | --- | --- |
Expand Down Expand Up @@ -95,6 +95,7 @@ Next available ID: `OPT-0089`.
| OPT-0086 | Planner-enabled downstream scheduling | The already integrated exact OPT-0080 DiT/VAE depth-two policies depend on authenticated graph/package topology, not on whether conditioning was produced by the planner, so removing the planner-only selector exclusion can recover the proven downstream scheduling saving without changing planner or model math | pending | benchmark-only | Registered before implementation; require selector-negative tests plus planner-enabled control/candidate final-latent, waveform, WAV, cancellation, and topology identity before widening production | [record](experiments/OPT-0086-planner-downstream-depth2-scheduling.md) | allocation `1ddb65e751529936b3ef3cd48a6360386c7dd205`; no implementation yet |
| OPT-0087 | Planner package-native dense GEMV | Route the already exact direct M1/M2 packed-BF16 kernel through all real planner decode layers and tied-head slices so the package can realize its useful bandwidth without OPT-0083's artificial submit-and-drain boundary around every isolated layer sample | positive | pending-integration | Actual Chrome preserved every full logit, appended K/V-cache evidence word, status, token, Philox cursor, topology, cancellation, and lifecycle result. Direct B won `16/16`, lowered every M1/M2 layer/head/model/complete median, reached `2.229363x` aggregate layer speedup, and reduced aggregate model-through-readback median `272.150 → 156.200 ms` (`115.950 ms`, projected `117.110 s/1,010`). Cleanup balanced `3,683/3,683` buffers and `133/133` maps with zero live resources. The fresh nominal launch passed; all later level-1/2 thermal observations are disclosed. Complete trajectory and planner-enabled product gates remain mandatory | [record](experiments/OPT-0087-planner-package-native-low-row-gemv.md), [result](results/OPT-0087/result.json) | allocation `84470d877ba69c6ef7821871f9bdb82d8c757747`; unchanged OPT-0083 direct kernel; implementation/harness `980623e0f9ec918f7a537b6a90e341adaad6d4b0`; pending integration gates |
| OPT-0088 | Portable device support | Every subgroup-dependent production owner (OPT-0032/0037 dense K4, OPT-0051 K7 row-reuse, OPT-0048 ConvTranspose K4, attention query8/quad-query) can gain a workgroup-memory counterpart consuming the unchanged hosted packages, selected by the existing execution-profile machinery, so adapters without `subgroups` (Safari, Firefox, iOS) run the production graph instead of failing `FEATURE_UNAVAILABLE`; compatibility experiment, bounded slowdown expected and reported, `shader-f16` stays fail-closed | pending | pending-integration | Portable dense/K7/ConvTranspose owners landed with test-enforced bit-identical arithmetic (byte-equal WGSL arithmetic sections, re-exported rev7/rev8 index math); attention routes to the existing portable oracle (reordered-rounding vs subgroup reduction). End-to-end masked-subgroups waveform and timing gates pending | [record](experiments/OPT-0088-portable-no-subgroup-production-path.md) | kernels c272b2d/c48c050/373e90a; selection wiring pending |
| OPT-0089 | DiT weight quantization | Weight-only symmetric int8 (per-32-K-block fp16 scales, round-to-nearest, clamp ±127) fake-quantization of all 264 rev7 DiT GEMM tensors, dequantized in place and run through the completely unchanged production graph, preserves end-to-end 30 s waveform quality within a small numerical envelope, so an int8-resident DiT (~1.51 GB + scales) is a credible answer to the observed iPhone 17 Safari OOM kill at `1,789,925,376 / 3,020,808,192` uploaded bytes (layer 14/24); pure quantization-damage gate, zero kernel changes, distinct mechanism from abandoned OPT-0058 activation-quantized DP4a | positive | benchmark-only | Per-tensor damage uniform and small: NRMSE `0.00515–0.00634` (median `0.00559`), min SNR `43.96 dB`, no outlier tensor/family, so no fp16-retention map needed. Determinism gate reproduced the pinned fp16 baseline WAV byte-exactly, then fake-quant vs fp16 on identical seeds gave lo-fi/12345 waveform NRMSE `0.0669` (Pearson `0.99777`, LSD ≈`3.5 dB`, RMS Δ `−0.050 dB`) and latin/424242 NRMSE `0.2268` (Pearson `0.97460`, LSD ≈`4.7 dB`, RMS Δ `+0.089 dB`; per-second max `5.19` is a near-silent-ending small-denominator artifact) — trajectory divergence of the 8-evaluation sampler, not noise-like corruption; zero non-finite samples and exact peak parity. Projected int8 DiT phase peak ≈`1.862 GB` tracked GPU (`1.51 GB` int8 + `94 MB` scales + fp16 norms/shared + measured `127 MB` overhead) versus the observed iPhone 17 kill at ≈`1.920 GB` — plausibly fits, marginal ≈`58 MB` margin; int8 kernel work justified, listening gate mandatory before any product claim | [record](experiments/OPT-0089-dit-int8-weight-fake-quant-gate.md), [quant result](results/OPT-0089/quant-error.json), [waveform result](results/OPT-0089/waveform-metrics.json) | `scripts/requantize-dit-int8.py` (repo root); fake-quant package `ef8355b9…` (models-local, not hosted); benchmark-only, no kernel or production change |

Experiment IDs are allocated before code changes, never reused, and never
removed from this table.
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,193 @@
# OPT-0089 — DiT int8 weight fake-quant quality gate (iPhone memory feasibility)

## Status

- Evidence: `positive` (numerical feasibility only; explicitly no listening
claim)
- Disposition: `benchmark-only` (this experiment ships no kernel and
integrates nothing; it produces a quality verdict and a memory projection
that gate any future int8-weight DiT kernel work)
- Date: 2026-08-24
- Author/agent: Claude agent for Alex-Wengg
- Risk class: `approximate` (weight values change; activations, kernels,
graph, sampler, and scheduling are untouched)

## Motivation (measured breadcrumb)

A real iPhone 17 Safari tab was OOM-killed during DiT dense-weight upload at
`1,789,925,376 / 3,020,808,192` bytes (layer 14/24). The fp16 rev7 dense DiT
package (`d3fc0020…`, 48 shards, `ACE_OPT_0009_DIT_MIXED_LAYER_BYTES =
3,020,808,192`) cannot fit an iPhone Safari tab alongside the rest of the
phase residency. An int8-weight path would roughly halve DiT resident weight
bytes, but before any int8 kernel is designed the pure quantization damage
must be measured with zero kernel changes.

## Hypothesis

Weight-only symmetric int8 quantization — one fp16 scale per 32-element block
along the input (K) dimension, round-to-nearest, clamp ±127 — applied to all
264 rank-two DiT GEMM/projection tensors in the rev7 package (216 fp16
`dit-gemm-n256-k32-tile-major-v1` + 48 packed-bf16
`dit-gemm-n128-k32-tile-major-v1` cross K/V projections; norms, biases,
scale-shift tables, and constants untouched) preserves end-to-end 30-second
generation quality within a small numerical envelope when the quantized
values are dequantized back to the original storage dtype and run through the
completely unchanged production graph (fake-quant). If true, an int8-resident
DiT (~1.51 GB weights + ~94 MB fp16 scales) is a credible iPhone path and
kernel work is justified; if quality collapses, int8-weight work stops here.

## Relationship to OPT-0058 (int8 DP4a, abandoned)

OPT-0058 quantized activations dynamically AND weights inside a DP4a compute
kernel and was abandoned when adversarial finite-to-zero collapse exceeded
its primitive gate before any timing. This experiment is a different
mechanism, not a revisit of that abandoned kernel: no activation
quantization, no DP4a, no kernel or arithmetic-order change of any kind —
only weight values move, and the damage is measured end-to-end on the real
product graph rather than on primitive adversarial fixtures. Any future int8
kernel remains subject to OPT-0058's recorded lessons.

## Method

1. `scripts/requantize-dit-int8.py` (python3 + numpy, streams shard-by-shard)
reads the verified local mirror of hosted package
`v1/dit-revision7/d3fc0020…`, fake-quantizes exactly the 264 GEMM tensors
in their native tile layout (a 32-K block is physically contiguous as
`[n_tile, k_block, :, n_in_tile]`), and writes a complete
content-addressed package tree. The manifest is rewritten textually so
only the 48 shard SHA-256 values and its own SHA-256 change; the manifest
byte length (254,357) and every tensor record stay identical, so the
unchanged runtime loads the package once its two pinned identity constants
(`DIT_MANIFEST_SHA256`, `ACE_OPT_0009_DIT_DENSE_MANIFEST_SHA256`) are
pointed at the new manifest SHA-256 for the run (temporary local patch,
never committed — the experiment is reproducible from the script plus the
documented invocation).
2. Per-tensor quantization error (RMSE, NRMSE, max abs, max relative vs
tensor amax, SNR, zero-scale blocks) is emitted to a JSON report and
persisted under `optimization/results/OPT-0089/`.
3. End-to-end: Chrome (puppeteer) against a dev server whose
`VITE_ACE_MODEL_ORIGIN` serves the local packages; fresh browser profile
per run; deterministic seeds.

## Gates

1. Environment determinism: a 30 s seed-12345 lo-fi fixture regenerated from
the UNMODIFIED local packages must reproduce the pinned known-good
baseline WAV SHA-256
`095267d7be0317ed9af10c64b8495d573b43207cd86615b6bbe66f27dc17895d`
before any fake-quant run is interpreted.
2. Quality comparison, fake-quant vs fp16 baseline, on two content-distinct
prompt/seed pairs (lo-fi hip hop seed 12345; the default Latin-percussion
product prompt seed 424242), 30 s each: waveform NRMSE, max |diff|,
per-1-second-segment NRMSE profile, STFT log-magnitude spectral distance,
and peak/RMS deltas.
3. Verdict recorded in the ledger: numerical envelope for both pairs,
per-tensor error outliers (with a mixed-precision fp16-retention map if
any tensor family is disproportionately damaged), and the projected
iPhone phase-peak residency versus the observed ~1.79 GB kill point.

This experiment makes no listening, timing, or product claim. The standard
listening-gate discipline applies before any int8 path could ever become a
product profile; numerical closeness here authorizes only kernel-design
work under a new ID.

## Identity

- Source package: local verified mirror of hosted `v1/dit-revision7/`
manifest `d3fc0020efcf60702db411da2fd4b93e9bb84f1437ed310aef01c892727e452f`
(source shard SHA-256s verified against the manifest before quantization)
- Fake-quant package: manifest
`ef8355b9cffff466b018b51275923982b071234933fe8a32897915eeeb01fa36`,
manifest byte length preserved at `254,357`; 264 GEMM tensors quantized
(`3,019,898,880` bytes), 193 passthrough tensors byte-identical
- Reference/VAE packages: unchanged hosted identities (`18f36c64…`,
`36a54d79…`) served from the same local origin
- Runtime: completely unchanged production graph and profiles
(`opt-0009-fp16-fp32-dense-v1` dense DiT,
`opt-0070-fixed32-quad-query32-full-self-production-v1` attention,
production VAE tuple); temporary local patch of the three pinned
`d3fc0020…` identity constants only, reverted after the runs
- Machine: Apple Silicon dev machine, stock Chrome via puppeteer, fresh
browser profile per run, dev server with
`VITE_ACE_MODEL_ORIGIN=/models/…` local origins
- Output WAV SHA-256s: fp16 lo-fi `095267d7…` (pinned baseline reproduced),
fp16 latin `bad98ef7ffb0512b4138241c596669ae6f05ee0e7dc6790a9998a5aaf23ed043`,
int8 lo-fi `d46597d5641fd8cbe0b36ff0dc5ccc65bf5af5d7717a87b3f5cf3757ca2dcc25`,
int8 latin `221dce8174f58251b4a38bab770d5be39194ac3cda46fcf34be2bbe6650fb88f`

## Results

Authoritative artifacts:
[quant-error.json](../results/OPT-0089/quant-error.json),
[waveform-metrics.json](../results/OPT-0089/waveform-metrics.json).

### Per-tensor quantization damage (weight level)

Uniform and small across all 264 tensors: NRMSE `0.00515–0.00634`
(median `0.00559`), minimum SNR `43.96 dB`, maximum per-element error
`0.394%` of tensor amax, zero zero-scale blocks. Family medians are tightly
clustered (o_proj `0.00535`, k/v/q_proj `0.00550–0.00558`, up/gate
`0.00566–0.00571`, down_proj `0.00574`; worst single tensor
`layers.23.self_attn.v_proj` at `0.00634`). No tensor or family is a
disproportionate outlier, so no mixed-precision fp16-retention map is
required at the weight level.

### Gate 1 — environment determinism

The fp16 run from the unmodified local packages reproduced the pinned
baseline WAV byte-exactly (`095267d7…`). PASS.

### Gate 2 — fake-quant vs fp16, two prompt/seed pairs (30 s each)

- Lo-fi seed 12345: waveform NRMSE `0.0669` (SNR `23.50 dB`), Pearson
`0.99777`, max |diff| `0.3867`, per-second NRMSE median `0.0484` /
max `0.3264`, mean log-spectral distance `0.1772` log10 units
(≈ `3.5 dB`), p95 `0.3640`, peak Δ `+0.0000` (shared −1 dBFS peak
normalization), RMS Δ `−0.050 dB`.
- Latin-percussion default prompt seed 424242: waveform NRMSE `0.2268`
(SNR `12.89 dB`), Pearson `0.97460`, max |diff| `0.9170`, per-second
NRMSE median `0.1173`, mean log-spectral distance `0.2343`
(≈ `4.7 dB`), p95 `0.4139`, peak Δ `+0.0000`, RMS Δ `+0.089 dB`. The
headline per-second maximum `5.19` is a small-denominator artifact of the
ending: at second 27 the baseline is essentially silent (segment RMS
`0.00061` vs `0.00322`) because the fake-quant realization sustains the
final decay slightly longer (second 26 RMS `0.0864` vs `0.0382`);
active-music seconds run `0.08–0.36`.

Interpretation: the damage manifests as trajectory divergence of the
8-evaluation sampler, not noise-like corruption — zero non-finite samples,
exact peak parity, RMS within `0.09 dB`, high global correlation, flat
per-second profiles except the latin ending decay. Waveform NRMSE therefore
overstates perceptual change, but the outputs are not near-bit-identical
realizations; the standard listening gate is mandatory before any product
use of an int8 path.

### Memory projection (iPhone)

Quantized GEMM storage: `1,509,949,440` B int8 + `94,371,840` B fp16
scales. DiT-phase resident weights become `1,735,340,288` B (including
`909,312` B fp16 layer norms/tables and the `130,109,696` B fp16
reference-shared tensors) versus `3,150,917,888` B today. Adding the
measured desktop 30 s DiT-phase non-weight overhead
(`3,277,864,192 − 3,150,917,888 = 126,946,304` B tracked arena/control)
projects an int8 phase peak of ≈ `1.862 GB` tracked GPU bytes. The observed
iPhone 17 kill point was ≈ `1.920 GB` tracked GPU (`1,789,925,376` dense
bytes uploaded + `130,109,696` reference-shared already resident), and the
pipeline destroys the conditioning phase before DiT upload begins, so that
kill already reflects minimal-residency ordering. The projected int8 peak
therefore sits ≈ `58 MB` under the observed kill boundary: plausibly fits,
but marginal, and untracked tab overhead (JS staging, page, wasm) consumes
an unknown share of the real Safari budget. Identified fallback levers:
quantize the `130 MB` reference-shared DiT tensors, and/or int4 for the MLP
family (`1.81 GB` of the fp16 bytes) which this gate's uniform error
profile makes the natural next candidate.

### Verdict

Weight-only int8 per-32-K-block quantization does not collapse ACE-Step
generation: the numerical feasibility gate passes and the projected DiT
phase residency crosses under the observed iPhone OOM boundary. int8-weight
kernel work (int8-resident storage with in-kernel dequant, new experiment
ID) is justified. Per-denoise-step latent comparison and the standard
listening gate are the mandatory next quality checkpoints; this record
makes no listening, timing, or product claim.
Loading
Loading