docs(fold): Metal's OCRBench flips point at the build, not drafting - #414
Merged
Merged
Conversation
Follows CUDA's review of #412 (5881649818). - Withdraws "the 12 flips are not attributable to the build". The one repeat of the build's code reproduced all 200 predictions of rows 0-199, although 999 of the deployed run's 1000 requests drafted. That points at the build. Two of its changes can move 31b-nvfp4's numerics on Metal, and they are not separated: the MLX pin move (d9add9d1 -> 59d600b5), and ADR 0039, which removed a float32 round trip from every dense nvfp4 linear (#312). gemma4's MLX model, sampler and drafting code differ from 0.34.0's only in package paths. - The heading reads "the same score as 0.34.0": 25 predictions and 12 verdicts differ. - The prose states 174/200 and 175/200. - The think-on sentence is attributed to this host's drafted-thinking measurement (5824799988) and softened. - A Runs: line names the logs, the drivers and the generators. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
|
Review of #414 at
Verified, with rename detection from
Follow-ups, none blocking:
ADR 0039's own "Metal is unaffected in behaviour" line is the CUDA host's to fix, in a separate dated amendment.
|
This was referenced Sep 29, 2026
glennneuber
added a commit
that referenced
this pull request
Sep 29, 2026
docs(fold): the #414 review's follow-ups for Metal's OCRBench section
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This follows up CUDA's review of #412 (5881649818), which arrived after #412 merged. It changes docs only: the Metal section of the fold record.
The correction. #412 said the 12 OCRBench flips against 0.34.0 "are not attributable to the build", because every drafted request is its own sample path. That was a reason, not a measurement, and the data points the other way:
speculative decode statslines in the scratch server's log.d9add9d1 → 59d600b5, which CUDA's review names;8a7ba949),QuantizedMatmulapplied every dense nvfp4 linear's global scale through a float32 round trip, with no Metal branch in it or in the loader. That round trip misses 17 of 31b's 191 vision scales by one ulp (Help wanted (MLX-CUDA): does the gemma4 vision encoder move across the v0.34 fold on a non-Metal backend? #312). ADR 0039's note "Metal is unaffected in behaviour" rests onGatherQMMalone.The record now reads: the flips balance, 6 each way, so the score does not move. Their cause is not isolated, and the evidence points at the build rather than at drafting.
The smaller items:
summarize_extbench.py's format, so that fix belongs in the generator.The #375 comment (5875111957) makes the same claim. A correction follows there once this is open.
check_source_paths.py --changed-since origin/mainand the name scan are clean.macbook-pro-m5-max-128GB/mlx-metal🤖 Generated with Claude Code