Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,8 +40,8 @@ the downloaded model.

The site uses ACE-Step 1.5 Turbo in direct mode. The optional planner is
available in the underlying runtime but disabled on the public music page.
Advanced settings includes a full-size, experimental INT8 quality preview for
listening comparisons; it does not yet provide a smaller download.
Advanced settings includes an experimental packed INT8 preview. Its DiT is
1.70 GB instead of 3.02 GB (43.7% smaller) for listening comparisons.
See the [ACE-Step README](packages/acestep/README.md) for implementation and
validation details.

Expand Down
12 changes: 7 additions & 5 deletions music.html
Original file line number Diff line number Diff line change
Expand Up @@ -219,12 +219,14 @@ <h1>ACE-Step 1.5</h1>
Model
<select id="model-variant" name="model-variant">
<option value="production" selected>Production (recommended)</option>
<option value="int8-quality-preview">INT8 quality preview (experimental)</option>
<option id="int8-model-option" value="int8-quality-preview" hidden>
Compressed INT8 preview (experimental)
</option>
</select>
<small>
The preview tests audible quantization effects but is not compressed.
It is also 5.75 GB; switching from a cached production model adds a
separate 3.02 GB DiT download. Use the same seed to compare.
<small id="int8-model-hint" hidden>
The preview uses a 1.70 GB packed INT8 DiT instead of the
3.02 GB production DiT (43.7% smaller). Switching models downloads
a separate DiT package. Use the same seed to compare quality.
</small>
</label>
</div>
Expand Down
5 changes: 3 additions & 2 deletions packages/acestep/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,8 +18,9 @@ have a portable fallback. Source-audio editing, cover generation, and the VAE
encoder are outside the current scope. Support for phones and other browsers
must be validated on the target device.

The FluidAudio music page uses direct generation and a **5.75 GB** model cache.
Its planner is disabled. The package's development demo also exposes the
The FluidAudio music page uses direct generation and a **5.75 GB** production
model cache. Its optional packed INT8 preview uses **4.43 GB**. The planner is
disabled. The package's development demo also exposes the
planner and uses a different reference manifest. The site's selected model
packages are defined in [config.ts](../../src/engines/musicgen-acestep/config.ts).

Expand Down
8 changes: 8 additions & 0 deletions packages/acestep/model/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,9 @@ manifest, verifies the result independently, and atomically installs it. No
notebook, manually edited weight, or unrecorded shell step is part of the
package recipe.

`repack_dit_int8.py` is the deterministic OPT-0091 derivative step. It accepts
only the exact revision-7 DiT package and emits the packed INT8 preview.

The structure deliberately follows `../parakeet.wgsl/model`, pinned at
Parakeet commit `7ee112738262a6f5a0efd2f150748a4087432fbb`. ACE-Step has a
larger staged graph, so its source contracts and phase-oriented shard plan are
Expand Down Expand Up @@ -55,6 +58,11 @@ uv run --frozen --project model --python 3.13 \
# the authenticated revision-7 package after complete staging verification.
uv run --frozen --project model --python 3.13 \
python3 model/convert.py --profile fp16-dit-dense-experimental --offline

# OPT-0091 packed INT8 preview from the authenticated revision-7 DiT package.
uv run --frozen --project model --python 3.13 \
python3 model/repack_dit_int8.py \
model/files-fp16-dit-rev7-oracle model/files-int8-dit-rev9
```

`--profile production` downloads and authenticates the pinned upstream files,
Expand Down
210 changes: 210 additions & 0 deletions packages/acestep/model/repack_dit_int8.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,210 @@
"""Create the authenticated OPT-0091 packed-INT8 DiT runtime package.

The input is the exact revision-7 mixed DiT package. Only the 216 repeated
dense matrices are repacked; cross-attention K/V and support tensors retain
their authenticated BF16 storage.
"""

from __future__ import annotations

import argparse
import hashlib
import json
import os
import tempfile
from pathlib import Path

import numpy as np

SOURCE_MANIFEST_SHA256 = (
"d3fc0020efcf60702db411da2fd4b93e9bb84f1437ed310aef01c892727e452f"
)
SOURCE_LAYOUT = "dit-gemm-n256-k32-tile-major-v1"
SOURCE_TRANSFORMATION = "bf16-to-ieee-fp16-dit-gemm-n256-k32-tile-major-v1"
PACKED_LAYOUT = "dit-gemm-n256-k32-int8-fp16-scale-tile-major-v1"
PACKED_TRANSFORMATION = (
"bf16-to-symmetric-int8-fp16-scale-n256-k32-tile-major-v1"
)
PACKED_DTYPE = "uint32-int8-fp16-blocks"
PACKED_WORDS_PER_TILE = 2_176
EXPECTED_DENSE_TENSORS = 216
EXPECTED_WEIGHT_BYTES = 1_699_602_432


def _canonical_json(value: object) -> bytes:
return (
json.dumps(value, ensure_ascii=False, separators=(",", ":"), sort_keys=True)
.encode("utf-8")
+ b"\n"
)


def _sha256(payload: bytes) -> str:
return hashlib.sha256(payload).hexdigest()


def _verify_runtime_files(directory: Path, manifest: dict[str, object]) -> None:
for record in manifest["files"]:
if record["kind"] != "weights":
continue
path = directory / record["name"]
payload = path.read_bytes()
if len(payload) != record["byteLength"] or _sha256(payload) != record["sha256"]:
raise ValueError(f"packed shard identity changed: {record['name']}")


def pack_dense_tensor(payload: bytes, shape: list[int]) -> tuple[bytes, list[int]]:
"""Pack one rev7 [N,K] FP16 tile-major matrix deterministically."""

if len(shape) != 2:
raise ValueError("packed INT8 dense tensor must be rank two")
columns, inner = shape
if columns % 256 != 0 or inner % 32 != 0:
raise ValueError("packed INT8 dense dimensions must divide N256/K32")
expected_bytes = columns * inner * 2
if len(payload) != expected_bytes:
raise ValueError("packed INT8 source tensor byte length changed")

tiles = np.frombuffer(payload, dtype="<f2").reshape(
columns // 256, inner // 32, 32, 256
)
output = bytearray()
for column_tile in range(columns // 256):
for inner_tile in range(inner // 32):
values = tiles[column_tile, inner_tile].astype(np.float32)
scales = (np.max(np.abs(values), axis=0) / np.float32(127.0)).astype(
"<f2"
)
divisors = scales.astype(np.float32)
divisors = np.where(divisors == 0.0, np.float32(1.0), divisors)
quantized = np.clip(
np.rint(values / divisors[None, :]), -127, 127
).astype(np.int8)
output.extend(quantized.tobytes(order="C"))
output.extend(scales.tobytes(order="C"))

storage_shape = [columns // 256, inner // 32, PACKED_WORDS_PER_TILE]
if len(output) != np.prod(storage_shape, dtype=np.int64) * 4:
raise ValueError("packed INT8 tensor storage shape changed")
return bytes(output), storage_shape


def repack(source: Path, output_root: Path) -> tuple[Path, dict[str, int | str]]:
source_manifest_bytes = (source / "manifest.json").read_bytes()
if _sha256(source_manifest_bytes) != SOURCE_MANIFEST_SHA256:
raise ValueError("OPT-0091 source manifest identity changed")
manifest = json.loads(source_manifest_bytes)
if (
manifest.get("profile") != "fp16-dit-dense-experimental"
or manifest.get("provenance", {}).get("converterRevision") != 7
):
raise ValueError("OPT-0091 requires the authenticated revision-7 package")

tensors_by_shard: dict[str, list[tuple[str, dict[str, object]]]] = {}
for name, tensor in manifest["tensors"].items():
tensors_by_shard.setdefault(tensor["shard"], []).append((name, tensor))

dense_count = 0
source_dense_bytes = 0
packed_dense_bytes = 0
new_files: list[dict[str, object]] = []
output_root.mkdir(parents=True, exist_ok=True)
with tempfile.TemporaryDirectory(prefix=".opt-0091-", dir=output_root) as temporary:
stage = Path(temporary)
for file_record in manifest["files"]:
if file_record["kind"] != "weights":
new_files.append(dict(file_record))
continue
source_path = source / file_record["name"]
source_bytes = source_path.read_bytes()
if (
len(source_bytes) != file_record["byteLength"]
or _sha256(source_bytes) != file_record["sha256"]
):
raise ValueError(f"source shard identity changed: {file_record['name']}")
output = bytearray()
tensors = sorted(
tensors_by_shard[file_record["name"]],
key=lambda item: item[1]["byteOffset"],
)
for name, tensor in tensors:
output.extend(b"\0" * (-len(output) % 256))
start = tensor["byteOffset"]
payload = source_bytes[start : start + tensor["byteLength"]]
tensor["byteOffset"] = len(output)
if tensor["layout"] == SOURCE_LAYOUT:
if (
tensor["dtype"] != "float16"
or tensor["transformation"] != SOURCE_TRANSFORMATION
):
raise ValueError(f"dense source contract changed: {name}")
source_dense_bytes += len(payload)
payload, storage_shape = pack_dense_tensor(
payload, tensor["logicalShape"]
)
tensor.update(
dtype=PACKED_DTYPE,
layout=PACKED_LAYOUT,
transformation=PACKED_TRANSFORMATION,
storageShape=storage_shape,
byteLength=len(payload),
)
dense_count += 1
packed_dense_bytes += len(payload)
output.extend(payload)
output.extend(b"\0" * (-len(output) % 256))
destination = stage / file_record["name"]
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_bytes(output)
new_files.append(
{
**file_record,
"byteLength": len(output),
"sha256": _sha256(output),
}
)

weight_bytes = sum(
record["byteLength"] for record in new_files if record["kind"] == "weights"
)
if dense_count != EXPECTED_DENSE_TENSORS or weight_bytes != EXPECTED_WEIGHT_BYTES:
raise ValueError("OPT-0091 packed inventory changed")
manifest["profile"] = "int8-dit-dense-experimental"
manifest["files"] = new_files
manifest["provenance"]["converterRevision"] = 9
manifest_bytes = _canonical_json(manifest)
manifest_sha256 = _sha256(manifest_bytes)
(stage / "manifest.json").write_bytes(manifest_bytes)
destination = output_root / manifest_sha256
if destination.is_symlink():
raise ValueError("OPT-0091 output must not be a symlink")
if destination.exists():
if (destination / "manifest.json").read_bytes() != manifest_bytes:
raise ValueError("existing OPT-0091 output does not match its digest")
_verify_runtime_files(destination, manifest)
else:
os.rename(stage, destination)
_verify_runtime_files(destination, manifest)

metrics: dict[str, int | str] = {
"manifestSha256": manifest_sha256,
"manifestBytes": len(manifest_bytes),
"denseTensorCount": dense_count,
"sourceDenseBytes": source_dense_bytes,
"packedDenseBytes": packed_dense_bytes,
"weightBytes": weight_bytes,
}
return destination, metrics


def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("source", type=Path)
parser.add_argument("output_root", type=Path)
args = parser.parse_args()
destination, metrics = repack(args.source.resolve(), args.output_root.resolve())
print(json.dumps({**metrics, "directory": str(destination)}, indent=2))


if __name__ == "__main__":
main()
40 changes: 40 additions & 0 deletions packages/acestep/model/tests/test_repack_dit_int8.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
import unittest
import sys
from pathlib import Path

import numpy as np

sys.path.insert(0, str(Path(__file__).resolve().parents[1]))

from repack_dit_int8 import PACKED_WORDS_PER_TILE, pack_dense_tensor


class RepackDitInt8Test(unittest.TestCase):
def test_packs_one_n256_k32_tile_with_fp16_scales(self) -> None:
source = np.arange(32 * 256, dtype=np.float32).reshape(32, 256)
source = ((source % 255) - 127).astype("<f2")

payload, shape = pack_dense_tensor(source.tobytes(), [256, 32])

self.assertEqual(shape, [1, 1, PACKED_WORDS_PER_TILE])
self.assertEqual(len(payload), 8_704)
quantized = np.frombuffer(payload[:8_192], dtype=np.int8).reshape(32, 256)
scales = np.frombuffer(payload[8_192:], dtype="<f2")
expected_scales = (
np.max(np.abs(source.astype(np.float32)), axis=0) / np.float32(127)
).astype("<f2")
np.testing.assert_array_equal(scales, expected_scales)
expected_quantized = np.clip(
np.rint(source.astype(np.float32) / expected_scales.astype(np.float32)),
-127,
127,
).astype(np.int8)
np.testing.assert_array_equal(quantized, expected_quantized)

def test_rejects_non_native_shape(self) -> None:
with self.assertRaisesRegex(ValueError, "N256/K32"):
pack_dense_tensor(bytes(255 * 32 * 2), [255, 32])


if __name__ == "__main__":
unittest.main()
3 changes: 2 additions & 1 deletion packages/acestep/optimization/LEDGER.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
Stage 2 was explicitly authorized on 2026-08-13. The approved baseline is
frozen and measured optimization is active.

Next available ID: `OPT-0091`.
Next available ID: `OPT-0092`.

| ID | Subsystem | Hypothesis | Evidence | Disposition | Result | Record | Implementation |
| --- | --- | --- | --- | --- | --- | --- | --- |
Expand Down Expand Up @@ -97,6 +97,7 @@ Next available ID: `OPT-0091`.
| OPT-0088 | Portable device support | Every subgroup-dependent production owner (OPT-0032/0037 dense K4, OPT-0051 K7 row-reuse, OPT-0048 ConvTranspose K4, attention query8/quad-query) can gain a workgroup-memory counterpart consuming the unchanged hosted packages, selected by the existing execution-profile machinery, so adapters without `subgroups` (Safari, Firefox, iOS) run the production graph instead of failing `FEATURE_UNAVAILABLE`; compatibility experiment, bounded slowdown expected and reported, `shader-f16` stays fail-closed | pending | pending-integration | Portable dense/K7/ConvTranspose owners landed with test-enforced bit-identical arithmetic (byte-equal WGSL arithmetic sections, re-exported rev7/rev8 index math); attention routes to the existing portable oracle (reordered-rounding vs subgroup reduction). End-to-end masked-subgroups waveform and timing gates pending | [record](experiments/OPT-0088-portable-no-subgroup-production-path.md) | kernels c272b2d/c48c050/373e90a; selection wiring pending |
| OPT-0089 | DiT weight quantization | Weight-only symmetric int8 (per-32-K-block fp16 scales, round-to-nearest, clamp ±127) fake-quantization of all 264 rev7 DiT GEMM tensors, dequantized in place and run through the completely unchanged production graph, preserves end-to-end 30 s waveform quality within a small numerical envelope, so an int8-resident DiT (~1.51 GB + scales) is a credible answer to the observed iPhone 17 Safari OOM kill at `1,789,925,376 / 3,020,808,192` uploaded bytes (layer 14/24); pure quantization-damage gate, zero kernel changes, distinct mechanism from abandoned OPT-0058 activation-quantized DP4a | positive | benchmark-only | Per-tensor damage uniform and small: NRMSE `0.00515–0.00634` (median `0.00559`), min SNR `43.96 dB`, no outlier tensor/family, so no fp16-retention map needed. Determinism gate reproduced the pinned fp16 baseline WAV byte-exactly, then fake-quant vs fp16 on identical seeds gave lo-fi/12345 waveform NRMSE `0.0669` (Pearson `0.99777`, LSD ≈`3.5 dB`, RMS Δ `−0.050 dB`) and latin/424242 NRMSE `0.2268` (Pearson `0.97460`, LSD ≈`4.7 dB`, RMS Δ `+0.089 dB`; per-second max `5.19` is a near-silent-ending small-denominator artifact) — trajectory divergence of the 8-evaluation sampler, not noise-like corruption; zero non-finite samples and exact peak parity. Projected int8 DiT phase peak ≈`1.862 GB` tracked GPU (`1.51 GB` int8 + `94 MB` scales + fp16 norms/shared + measured `127 MB` overhead) versus the observed iPhone 17 kill at ≈`1.920 GB` — plausibly fits, marginal ≈`58 MB` margin; int8 kernel work justified, listening gate mandatory before any product claim | [record](experiments/OPT-0089-dit-int8-weight-fake-quant-gate.md), [quant result](results/OPT-0089/quant-error.json), [waveform result](results/OPT-0089/waveform-metrics.json) | `scripts/requantize-dit-int8.py` (repo root); fake-quant package `ef8355b9…` (models-local, not hosted); benchmark-only, no kernel or production change |
| OPT-0090 | DiT quantization listening preview | The authenticated OPT-0089 fake-quant package can be exposed as an explicit opt-in listening comparison without changing the production default, weakening package identity checks, or claiming packed-int8 size/runtime benefits | positive | integrated | Reproduced the recorded `ef8355b9…` package and byte-identical quant report, published the immutable full-size artifact, then added an Advanced listening selector with orderly model switching and explicit size/non-production disclosure. 43 web integration, 4 unit, and 2,029 ACE tests plus typecheck, formatting, production build, and all PR checks passed; fresh browser/GPU listening remains external | [record](experiments/OPT-0090-int8-quality-preview.md) | implementation `2816d6b83185cbde5abd6780a9f5ac903a655154`; artifact `ac50b5c854fb044ce058acb91d4cd9ab82d99cfa` |
| OPT-0091 | Packed weight-only INT8 DiT runtime | Keep the 216 repeated-layer dense matrices as signed INT8 with per-output/K32 FP16 scales and dequantize inside subgroup and portable kernels, preserving the current activation rounding and increasing-K FP32 accumulation while materially reducing the 3.02 GB DiT package | pending | experimental-preview | Deterministic packed package is 1.70 GB (43.7% smaller); manifest, hosted identity, automated tests, and build passed. Actual Chrome/WebGPU, waveform, and listening gates remain external, so production stays default and no quality/mobile claim is made | [record](experiments/OPT-0091-packed-int8-dit-runtime.md) | artifact `a3233c9f…`; HF `bc43ba20409825c13d7ef25694d39ac47dd8c9a4` |

Experiment IDs are allocated before code changes, never reused, and never
removed from this table.
Expand Down
Loading
Loading