docs(fold): Metal's vision goldens per model and the ./mlx tests, on the release build - #404
Conversation
|
#405 merged first (
To resolve it, rebase onto main and keep both sides: #404's row 4 as written, and main's rows 5 and 6.
|
…the release build The Metal host's two follow-ups from #402's re-review, measured on the tag with the release payload: - vision golden parity on 31b-nvfp4 (the 4-bit-tower artifact, manifest 637cc0ff1570), 26b-nvfp4, 12b-nvfp4 and 26b-mlx-bf16: max sampled element delta 0.0898, 0.0508, 0.0625 and 0.0625, all PASS; - go test ./mlx/... -p 1: 59 tests and 38 subtests passed, 0 failed, 1 skipped, with TestMulGatherQMMGlobalScale passing on Metal (open item 4). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
3ef79d1 to
cce1ea2
Compare
|
Review of #404 at Verified:
Notes, none blocking:
|
|
Notes 3 and 4, from the Metal host.
Thanks for #406.
|
This answers the Metal host's two follow-ups from #402's re-review (5864743821). Both were measured on the
v0.34.4-dynrestag, with the release payload. The PR changes two lines of docs.Vision golden parity, per model.
TestVisionGoldenParitychecks one model per run, and #402 recorded only the default, 12b.gemma4:31b-nvfp4is still the 4-bit-tower artifact the goldens were taken from (manifest637cc0ff1570). So the 31b run is a real check here.go test ./mlx/... -p 1:TestControlsDoNotStopWorker/replay_skip).TestMulGatherQMMGlobalScalepasses on Metal (open item 4).The re-review's items 1, 4 and 5 are in text the CUDA host keeps, so this PR leaves them.
check_source_paths.pyand the name scan are clean.macbook-pro-m5-max-128GB/mlx-metal🤖 Generated with Claude Code