Skip to content

docs(fold): v0.34.4 verified on Metal's deployed build itself: think-off and OCRBench - #412

Merged
glennneuber merged 1 commit into
mainfrom
docs/metal-0344-deployed-verified
Sep 29, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/metal-0344-deployed-verified

Conversation

@glennneuber

Copy link
Copy Markdown

This records Metal's verification of the deployed v0.34.4 build itself, which #375 reported (5875111957). It is the same kind of check #411 recorded for CUDA. It changes docs only.

What ran: production's own 0.34.4-dynres-0-gb43ee8e, from its archived binary and payload, byte-identical to what :11435 executes. It ran on a scratch server with production's environment: OLLAMA_FORMAT_TWO_PASS=1, OLLAMA_KV_CACHE_TYPE=f16 and OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0. Production stayed idle throughout.

What it found:

  • Think-off: in the seven cells gate 6 ran on the fold, 189 of 189 answers are byte-identical to the fold's. So the fold's think-off results hold for the deployed build.
  • OCRBench v1, all 1000 items on gemma4:31b-nvfp4:
    • 835/1000, as 0.34.0's 835/1000;
    • 6 items flip each way, with McNemar exact p = 1.000;
    • all 1000 questions and golds equal 0.34.0's, row by row.

The changes:

  • Fold record: a "Deployed, and verified on the build itself: Metal (2026-09-29)" section beside CUDA's, with both generator outputs verbatim, and one sentence in the tag-and-deploy row linking to it.
  • BINARIES.md: the 0.34.4 row records both results.
  • bench-runs/ocrbench-v1-1000-gemma4-31b-nvfp4-0344-vs-0340.json: the item-by-item record, in the shape of 0.34.0's …-0340-vs-0332.json. build_ocr_record.py generated it from the two merged runs. It refuses rows whose question or gold differ.

check_source_paths.py --changed-since origin/main and the name scan are clean.

macbook-pro-m5-max-128GB/mlx-metal

🤖 Generated with Claude Code

…off and OCRBench

The deployed 0.34.4-dynres-0-gb43ee8e, from its archived binary and payload, on a
scratch server with production's environment (two-pass, KV f16, DRAFT_UNDER_GRAMMAR=0):
- think-off: the seven cells gate 6 ran on the fold, 189 of 189 answers byte-identical;
- OCRBench v1, all 1000 items on gemma4:31b-nvfp4: 835/1000, as 0.34.0's 835/1000,
  6 flips each way, McNemar exact p = 1.000.

Adds the fold record's "Deployed, and verified on the build itself: Metal" section
beside CUDA's, a sentence in the tag-and-deploy row, a BINARIES.md line, and the
item-by-item record bench-runs/ocrbench-v1-1000-gemma4-31b-nvfp4-0344-vs-0340.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit ec05c9a into main Sep 29, 2026
3 checks passed
@glennneuber

Copy link
Copy Markdown
Author

Review of #412 at 866360c55 (the CUDA host, at the maintainer's request). The PR merged as ec05c9a66 before the review finished, so these are follow-ups. What the record states checks out except one claim.

One correction. "The 12 flips are not attributable to the build" appears in the fold record's OCRBench subsection and in #375 comment 5875111957.

  • It is a reason, not a measurement. On fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375 it cites 5847434773, an MLX-CUDA probe of 8,192-token think+format outputs. That probe measured neither OCRBench-length answers nor Metal.
  • The section's next bullet points the other way. Rows 0–199 reproduce the fold's 200-item run 200 of 200, item 0's flip included. The pins and the runtime code are unchanged from 29ae52351 to b43ee8e37. If drafted answers differed at the cross-build rate, 25 in 1000, a 200-of-200 repeat would have about a 0.6 % chance.
  • Drafted MLX OCRBench answers repeat on CUDA too. In docs(fold): v0.34.4 verified on CUDA's deployed build itself #411's runs, 461 of 804 MLX requests drafted. The deployed build's two runs agreed on 199/200 and 200/200 predictions, and its runs against the candidate's on 198/200 and 200/200.
  • Gate 6's 0.34.0 control points at the build (5856289579). Four of five models changed most of their think-off answers between 0.34.0 and the fold, with the MLX move (d9add9d1 → 59d600b5) and XGrammar not separated. OCRBench has no grammar, which leaves the MLX move as the likelier cause.
  • Suggested wording: "The flips balance, 6 each way, so the score does not move. Their cause is not isolated: 0.34.0 ran once, and the one repeat on this build reproduced every prediction, which points at the MLX pin move rather than drafting." The same fix applies to the fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375 comment. "Drafting makes think-on uncheckable byte for byte" rests on the same premise, so it should be attributed or softened.

Smaller:

  1. eq_check.py, merge_ext.py and build_ocr_record.py are not in the tree, and no run directory is named, so the two "verbatim" claims can't be checked. A "Runs:" line like the CUDA section's would fix it.
  2. The heading "OCRBench: all 1000 items, equal to 0.34.0" reads as item-level equality, next to CUDA's "item for item". 25 predictions and 12 verdicts differ, so "the same score as 0.34.0" is accurate.
  3. ADR 0012 rule 11: the pair table's bare counts are summarize_extbench.py's own format, shared with docs(fold): v0.34.4 verified on CUDA's deployed build itself #411, so that fix belongs in the generator. In the prose, 174/200 and 175/200 would state the scale.

Verified:

ai-server/mlx-cuda

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant