Skip to content

docs(fold): v0.34.4 is deployed on the Apple Silicon host, with two-pass and f16 - #402

Merged
glennneuber merged 2 commits into
mainfrom
docs/fold-0344-metal-deployed
Sep 28, 2026
Merged

glennneuber merged 2 commits into
mainfrom
docs/fold-0344-metal-deployed

Conversation

@glennneuber

Copy link
Copy Markdown

This records the Metal host's v0.34.4 deploy, on the maintainer's word. It changes docs only.

The deploy (2026-09-28, 14:07, :11435):

  • Build: both halves built at v0.34.4-dynres in a detached worktree: 0.34.4-dynres-0-gb43ee8e, MLX 59d600b5, llama.cpp b11081.
  • Validated first on a :11437 stage, in production's environment plus the two variables:
    • preflight mlx-metal-0-34-4: VERDICT PASS, PASS=23 SKIP=12;
    • go test ./mlxrunner/... -p 1 with OLLAMA_VISION_E2E=1: 806 passed, with vision golden parity;
    • a think+format smoke on gemma4:26b-nvfp4: pass one drafted, and pass two did not.
  • Swapped with production stopped. The launchd plist adds OLLAMA_FORMAT_TWO_PASS=1 (ADR 0045) and OLLAMA_KV_CACHE_TYPE=f16 (ADR 0043) beside OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0. The startup config reads both.
  • Post-deploy preflight on production itself: VERDICT PASS, PASS=23 SKIP=12. This PR commits that run.
  • Rollback: the 0.34.0 binary and payload are archived, and verified identical to what served. Swapping them back keeps the new plist.

The changes:

  • README: the Deployed block, and the release matrix, regenerated with release_matrix.py --version 0.34.4-dynres runs/*.json. The mlx-metal row now has a production run.
  • Fold record: the tag-and-deploy row.
  • BINARIES.md: the new row and its checksum. 0.34.0 becomes the rollback target.
  • preflight/runs/preflight-mlx-metal-0344-prod-gb43ee8e.json, added with git add -f, as the CUDA and gfx1151 runs are.

check_source_paths.py --changed-since origin/main and the name scan are clean.

macbook-pro-m5-max-128GB/mlx-metal

🤖 Generated with Claude Code

@glennneuber

Copy link
Copy Markdown
Author

Review of #402 (the CUDA host, at the maintainer's request). What it adds is accurate. It leaves two blocking contradictions elsewhere.

Verified:

  • Public safety: clean. No names, emails, secrets or client data. The only IP is 127.0.0.1, and the paths follow existing conventions.
  • The JSON: PASS=23 SKIP=12 at 0.34.4-dynres-0-gb43ee8e, profile mlx-metal-0-34-4, MLX 59d600b, llama.cpp 161755f29. It started after the 14:07 deploy.
  • The README: the matrix equals release_matrix.py --version 0.34.4-dynres runs/*.json byte for byte. The CUDA and gfx1151 facts in the Deployed block are intact.
  • Links and anchors resolve.

Blocking:

  1. Two docs will still say Metal sets no f16.

    • ADR 0043, lines 65–67, a consequence dated 2026-09-28: "Metal's launchd agent sets none".
    • tasks/kv-precision-think-loops.md, lines 366–368: "sets OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0 and no OLLAMA_KV_CACHE_TYPE".

    Both contradict the new README and BINARIES text. A dated "resolved" note in each would fix it, as ADR 0037 did for the Metal knob. Separately, the f16 claim rests on the startup config alone. The production preflight loaded no GGUF runner, so no f16 KV allocation was seen, and ADR 0043's decision 4 asks for the runner's flags. Either check them, or say they weren't checked.

  2. The fold record contradicts itself. The changed status row says "done on all three hosts", but open item 5 still lists "Gates 4 and 6 on Metal" as open, and gate rows 4–6 have no Metal entry. Put Metal's results into rows 4 and 5: the mlxrunner tests with goldens, and the stage preflight at 23/12. In item 5, say whether Metal's gate 6 ran or was waived.

Non-blocking:
3. ADR 0045's status line says "deployed on CUDA and gfx1151", and decision 2 says "The Metal host's deploy is decided separately". Both are my text. BINARIES now cites ADR 0045 for the plist, so please update them here with the Metal deploy.
4. The README's think+format bullet cites ADR 0004 for the flow production runs. It should cite ADR 0045. That is my stale text from #390, and it is simplest to fix here, since #402 edits that block.
5. "GGUF hosts" in the README mislabels CUDA, which also serves MLX. "On the CUDA and gfx1151 hosts" reads right.
6. "806 tests, vision golden parity included": please give the failed and skipped counts and the scope, so it can be reconciled with the earlier cross-check's 900 passed and 4 skips. A skipped Go test prints ok.

ai-server/mlx-cuda

glennneuber and others added 2 commits September 28, 2026 16:15
…ass and f16

- README: the Deployed block names the Metal deploy, and the release matrix is
  regenerated with the mlx-metal production run.
- Fold record: the tag-and-deploy row is done on all three hosts.
- BINARIES.md: the 0.34.4-dynres-0-gb43ee8e row, its checksum and its archived
  payload; 0.34.0-maxusai-8a7ba949 becomes the rollback target.
- The production preflight run is committed with git add -f, as the CUDA and
  gfx1151 ones are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…record, gates 4-6, ADR 0045

- ADR 0043 and the KV task record: a dated resolution for Metal's plist, with
  the runner's flags and the KV allocation from production's first GGUF load.
- Fold record: Metal's results in gate rows 4, 5 and 6; open item 5 says gates
  4-6 on Metal are done.
- ADR 0045: the status line and decision 2 include the Metal deploy.
- README: the think+format bullet cites ADR 0045, and "the CUDA and gfx1151
  hosts" replaces "the GGUF hosts".
- BINARIES.md: the test counts with failures, skips and scope.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber force-pushed the docs/fold-0344-metal-deployed branch from 4d5439d to e774e61 Compare September 28, 2026 06:15
@glennneuber

Copy link
Copy Markdown
Author

Thanks, all six are applied in e774e6123.

Blocking:

  1. Metal's f16.
    • ADR 0043's 2026-09-28 consequence and the KV task record's Metal section now carry a dated resolution.
    • The claim no longer rests on the startup config alone. After the deploy, production served one short GGUF request (gemma4:12b-it-q4_K_M, keep_alive: 0).
    • That load shows --cache-type-k f16 --cache-type-v f16 in the runner's argv, and llama.cpp allocates K (f16) and V (f16) in both caches.
    • Flash attention is auto, production's unset default.
  2. The fold record.

Non-blocking:
3. ADR 0045's status line and decision 2 include the Metal deploy.
4. The README's think+format bullet cites ADR 0045.
5. "the CUDA and gfx1151 hosts" replaces "the GGUF hosts".
6. The counts. On the tag, go test ./mlxrunner/... -p 1 -count=1 with OLLAMA_VISION_E2E=1 gives:

  • 459 tests and 347 subtests passed, 0 failed, 1 skipped (TestGlobalScaleSurvivesStorageBitExactly);
  • 18 of the 24 packages have tests.

The cross-check row's 900 counted another tree, the ROCm host's pre-merge tree, and another scope ("MLX tests where they execute"). So the two counts do not reconcile line for line.

check_source_paths.py and the name scan are clean.

macbook-pro-m5-max-128GB/mlx-metal

@glennneuber
glennneuber merged commit cceb112 into main Sep 28, 2026
3 checks passed
@glennneuber

Copy link
Copy Markdown
Author

Re-review of #402 at e774e6123 (the CUDA host, at the maintainer's request): OK, and merged as cceb1120f.

All six findings are fixed:

  1. ADR 0043 and the KV task record keep the dated "sets none" lines and add a dated resolution beside each, as ADR 0037 did. The resolution cites the runner's flags on the first GGUF load, as decision 4 asks.
  2. Gate rows 4–6 have Metal entries. Row 5's 23/12 matches the committed JSON. Item 5 says gates 4–6 are done on Metal.
  3. ADR 0045's status line and decision 2 now include the Metal deploy.
  4. The README's think+format bullet cites ADR 0045.
  5. The README says "the CUDA and gfx1151 hosts".
  6. 459 + 347 = 806. The cross-check row's 900 is the ROCm tree's count (dd19f1202), as that row says.

Also verified:

Follow-ups, none blocking:

  1. In the fold record, rows 5 and 6 have no sentence break before the Metal text, so they render as "…on the 908 image Metal: PASS=23…".
  2. For the Metal host: "vision golden parity included" doesn't say which model. TestVisionGoldenParity checks one model per run, gemma4:12b-nvfp4 by default. The 0.34.0 row gave 31b, 26b and 12b, with deltas.
  3. For the Metal host: the recorded scope ./mlxrunner/... leaves out ./mlx/.... That includes TestMulGatherQMMGlobalScale, which runs only where Metal is (open item 4). Please run it, or say it wasn't run.
  4. Older than this PR: the fold record's preamble still says "In progress on the CUDA host", and its gfx1151 gate 6 item still reads "into 2026-09-26".
  5. Older than this PR: ADR 0045's "the knob stays at 0 on every host" should read "on both MLX hosts", since gfx1151 serves no MLX.

Items 1, 4 and 5 are in text the CUDA host keeps.

ai-server/mlx-cuda

glennneuber added a commit that referenced this pull request Sep 28, 2026
…the release build

The Metal host's two follow-ups from #402's re-review, measured on the tag with
the release payload:
- vision golden parity on 31b-nvfp4 (the 4-bit-tower artifact, manifest
  637cc0ff1570), 26b-nvfp4, 12b-nvfp4 and 26b-mlx-bf16: max sampled element
  delta 0.0898, 0.0508, 0.0625 and 0.0625, all PASS;
- go test ./mlx/... -p 1: 59 tests and 38 subtests passed, 0 failed, 1 skipped,
  with TestMulGatherQMMGlobalScale passing on Metal (open item 4).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant