Skip to content

docs(adr): ADR 0039's change reaches Metal too, through QuantizedMatmul - #416

Merged
glennneuber merged 1 commit into
mainfrom
docs/adr0039-metal-amendment
Sep 29, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/adr0039-metal-amendment

Conversation

@glennneuber

Copy link
Copy Markdown

This adds a dated amendment to ADR 0039: its change reaches Metal too, through QuantizedMatmul. Docs only, one file.

The sentence it corrects. ADR 0039's consequences say "Metal is unaffected in behaviour: its GatherQMM still receives m × 2688". They add that Metal's 26b and 31b encoder numbers move only for mlx#3912, and that "this change does not interact with that". The Metal host pointed out in #414 that the claim rests on GatherQMM alone. The ADR's own wrapper table says otherwise for the dense path.

Checked at 0.34.0 (8a7ba949):

  • The loader stored f32(m) × 2688, and scaleAndCast divided by 2688 in float32.
  • Only GatherQMM takes a Metal branch. Among the scale wrappers, the only MetalIsAvailable branch is GatherQMM's (x/mlxrunner/mlx/ops_extra.go:137). The loader has none.
  • The other MetalIsAvailable hits are a benchmark, the runner and gemma4 model code. The vision tower's only makes the attention output contiguous, so its dense linears take the same path on every platform.
  • So QuantizedMatmul, which carries every dense nvfp4 linear, took the round trip on Metal too, and ADR 0039 removed it there.

It was measured. ADR 0037, which is about Metal, records 31b's golden delta going 0.1406 → 0.1094 with mlx#3912. With the round trip removed it reached 0.0898, which is MLX-CUDA's value exactly. Metal's goldens on the v0.34.4 release read 0.0898 (#404).

What changes:

  • The original sentence stays.
  • A bolded, dated amendment goes under it, as ADR 0012 and ADR 0043 do.
  • The status line gets a one-line note, as ADR 0033's does.
  • No other document repeats the claim.

Checks:

  • check_source_paths.py --changed-since origin/main is clean.
  • The name scan is clean.
  • The relative link to ADR 0037 resolves.

ai-server/mlx-cuda

🤖 Generated with Claude Code

ADR 0039's consequences said Metal is unaffected in behaviour because its GatherQMM still receives m x 2688, and
that Metal's 26b and 31b encoder numbers move only for mlx#3912. The ADR's own table says otherwise for the dense
path: of its three wrappers only GatherQMM branches on Metal, and QuantizedMatmul applies the scale itself on
every platform because mlx_quantized_matmul has no global-scale argument. So Metal's dense nvfp4 linears took the
round trip, and this change removes it there too. ADR 0037 measured it on Metal's 31b encoder: 0.1406 -> 0.1094
from mlx#3912, then 0.0898 with the round trip removed, MLX-CUDA's value exactly; Metal's v0.34.4 goldens read
0.0898 (#404). The original sentence stays, with a dated amendment under it and a note on the status line. The
Metal host found it (#414).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant