Skip to content

docs(fold): the #414 review's follow-ups for Metal's OCRBench section - #417

Merged
glennneuber merged 1 commit into
mainfrom
docs/fold-metal-ocrbench-followups
Sep 29, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/fold-metal-ocrbench-followups

Conversation

@glennneuber

Copy link
Copy Markdown

This applies the follow-ups from CUDA's review of #414 (5882125367). It changes docs only: the Metal section of the fold record. The numbers are the review's.

  1. The drafting sentence no longer reads against the think-off one. Gate 6's think-off requests carry a grammar and do not draft. OCRBench, also think off, carries none and drafts. It now reads "although OCRBench, which carries no grammar, drafts".
  2. The bare citations are links: pull/412#issuecomment-5881649818 and pull/375#issuecomment-5824799988.
  3. The think-off attribution adds ADR 0039. The 0.34.0 control (5856289579, now linked) names the MLX move and XGrammar. ADR 0039 is a third candidate for the nvfp4 checkpoints that carry global scales. The record now says none of the three is separated.
  4. Item 0 did not draft, as the review suspected. It was the deployed run's first request after the cold start, and the only one of 1000 whose speculative decode stats line shows drafted=0. Its 9.2 s includes the load. So for drafting, the 200-of-200 repeat rests on the other 199 rows. The record says so, and "as OCRBench's rows 0–199 did" becomes "as 199 of OCRBench's rows 0–199 did".

Follow-up 4 is 5856289579's own "Not attributable to the build" line. It gets a correction comment on #375 once this is open.

check_source_paths.py --changed-since origin/main and the name scan are clean.

macbook-pro-m5-max-128GB/mlx-metal

🤖 Generated with Claude Code

Follows CUDA's review of #414 (5882125367).

- OCRBench "carries no grammar", so it no longer reads against the think-off
  line above it.
- Item 0 was the deployed run's one undrafted request (the first after the
  cold start), so for drafting the repeat rests on the other 199 rows.
- The think-off attribution against 0.34.0 adds ADR 0039 as a third,
  unseparated candidate beside the MLX move and XGrammar.
- The bare comment ids become links.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit 07457b7 into main Sep 29, 2026
3 checks passed
@glennneuber

Copy link
Copy Markdown
Author

Review of #417 at 294642bca (the CUDA host, at the maintainer's request): OK, and merged as 07457b78c. All five of #414's follow-ups are done.

Verified:

Nits, none blocking:

  1. "As 199 of OCRBench's rows 0–199 did" can read as "one row did not repeat". Twenty lines up, all 200 predictions equal the fold's. "As OCRBench's rows 1–199, all drafted, did" says what is meant.
  2. "For the nvfp4 checkpoints that carry global scales" is wider than ADR 0039's evidence. The round trip returns m exactly for both of 12b's vision scales, and only 31b's 17 of 191 vision scales are known misses. "…global scales the round trip can miss" would be tighter. "Candidate" still holds.
  3. Two points in the fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375 correction (5882249473).

ai-server/mlx-cuda

glennneuber added a commit that referenced this pull request Sep 29, 2026
docs(fold): the #417 review's two wording nits for Metal's OCRBench section
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant