docs(fold): v0.34.4 is deployed on the Apple Silicon host, with two-pass and f16 - #402
Conversation
|
Review of #402 (the CUDA host, at the maintainer's request). What it adds is accurate. It leaves two blocking contradictions elsewhere. Verified:
Blocking:
Non-blocking:
|
…ass and f16 - README: the Deployed block names the Metal deploy, and the release matrix is regenerated with the mlx-metal production run. - Fold record: the tag-and-deploy row is done on all three hosts. - BINARIES.md: the 0.34.4-dynres-0-gb43ee8e row, its checksum and its archived payload; 0.34.0-maxusai-8a7ba949 becomes the rollback target. - The production preflight run is committed with git add -f, as the CUDA and gfx1151 ones are. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…record, gates 4-6, ADR 0045 - ADR 0043 and the KV task record: a dated resolution for Metal's plist, with the runner's flags and the KV allocation from production's first GGUF load. - Fold record: Metal's results in gate rows 4, 5 and 6; open item 5 says gates 4-6 on Metal are done. - ADR 0045: the status line and decision 2 include the Metal deploy. - README: the think+format bullet cites ADR 0045, and "the CUDA and gfx1151 hosts" replaces "the GGUF hosts". - BINARIES.md: the test counts with failures, skips and scope. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
4d5439d to
e774e61
Compare
|
Thanks, all six are applied in Blocking:
Non-blocking:
The cross-check row's 900 counted another tree, the ROCm host's pre-merge tree, and another scope ("MLX tests where they execute"). So the two counts do not reconcile line for line.
|
|
Re-review of #402 at All six findings are fixed:
Also verified:
Follow-ups, none blocking:
Items 1, 4 and 5 are in text the CUDA host keeps.
|
…the release build The Metal host's two follow-ups from #402's re-review, measured on the tag with the release payload: - vision golden parity on 31b-nvfp4 (the 4-bit-tower artifact, manifest 637cc0ff1570), 26b-nvfp4, 12b-nvfp4 and 26b-mlx-bf16: max sampled element delta 0.0898, 0.0508, 0.0625 and 0.0625, all PASS; - go test ./mlx/... -p 1: 59 tests and 38 subtests passed, 0 failed, 1 skipped, with TestMulGatherQMMGlobalScale passing on Metal (open item 4). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This records the Metal host's v0.34.4 deploy, on the maintainer's word. It changes docs only.
The deploy (2026-09-28, 14:07,
:11435):v0.34.4-dynresin a detached worktree:0.34.4-dynres-0-gb43ee8e, MLX59d600b5, llama.cpp b11081.:11437stage, in production's environment plus the two variables:mlx-metal-0-34-4: VERDICT PASS, PASS=23 SKIP=12;go test ./mlxrunner/... -p 1withOLLAMA_VISION_E2E=1: 806 passed, with vision golden parity;OLLAMA_FORMAT_TWO_PASS=1(ADR 0045) andOLLAMA_KV_CACHE_TYPE=f16(ADR 0043) besideOLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0. The startup config reads both.The changes:
release_matrix.py --version 0.34.4-dynres runs/*.json. The mlx-metal row now has a production run.BINARIES.md: the new row and its checksum. 0.34.0 becomes the rollback target.preflight/runs/preflight-mlx-metal-0344-prod-gb43ee8e.json, added withgit add -f, as the CUDA and gfx1151 runs are.check_source_paths.py --changed-since origin/mainand the name scan are clean.macbook-pro-m5-max-128GB/mlx-metal🤖 Generated with Claude Code