docs(adr): ADR 0039's change reaches Metal too, through QuantizedMatmul - #416
Merged
Merged
Conversation
ADR 0039's consequences said Metal is unaffected in behaviour because its GatherQMM still receives m x 2688, and that Metal's 26b and 31b encoder numbers move only for mlx#3912. The ADR's own table says otherwise for the dense path: of its three wrappers only GatherQMM branches on Metal, and QuantizedMatmul applies the scale itself on every platform because mlx_quantized_matmul has no global-scale argument. So Metal's dense nvfp4 linears took the round trip, and this change removes it there too. ADR 0037 measured it on Metal's 31b encoder: 0.1406 -> 0.1094 from mlx#3912, then 0.0898 with the round trip removed, MLX-CUDA's value exactly; Metal's v0.34.4 goldens read 0.0898 (#404). The original sentence stays, with a dated amendment under it and a note on the status line. The Metal host found it (#414). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds a dated amendment to ADR 0039: its change reaches Metal too, through
QuantizedMatmul. Docs only, one file.The sentence it corrects. ADR 0039's consequences say "Metal is unaffected in behaviour: its
GatherQMMstill receivesm × 2688". They add that Metal's 26b and 31b encoder numbers move only for mlx#3912, and that "this change does not interact with that". The Metal host pointed out in #414 that the claim rests onGatherQMMalone. The ADR's own wrapper table says otherwise for the dense path.Checked at 0.34.0 (
8a7ba949):scaleAndCastdivided by 2688 in float32.GatherQMMtakes a Metal branch. Among the scale wrappers, the onlyMetalIsAvailablebranch isGatherQMM's (x/mlxrunner/mlx/ops_extra.go:137). The loader has none.MetalIsAvailablehits are a benchmark, the runner and gemma4 model code. The vision tower's only makes the attention output contiguous, so its dense linears take the same path on every platform.QuantizedMatmul, which carries every dense nvfp4 linear, took the round trip on Metal too, and ADR 0039 removed it there.It was measured. ADR 0037, which is about Metal, records 31b's golden delta going
0.1406 → 0.1094with mlx#3912. With the round trip removed it reached0.0898, which is MLX-CUDA's value exactly. Metal's goldens on the v0.34.4 release read0.0898(#404).What changes:
Checks:
check_source_paths.py --changed-since origin/mainis clean.ai-server/mlx-cuda🤖 Generated with Claude Code