Skip to content

docs(fold): v0.34.4 verified on CUDA's deployed build itself - #411

Merged
glennneuber merged 1 commit into
mainfrom
docs/fold-0344-postverify-cuda
Sep 29, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/fold-0344-postverify-cuda

Conversation

@glennneuber

Copy link
Copy Markdown

This PR adds a section to the v0.34.4 fold record: "Deployed, and verified on the build itself: CUDA (2026-09-29)". The status table's "tag and deploy" row gets one sentence linking to it. It follows 0.34.2's record, which re-measured the shipped image after its deploy. Docs only, one file.

Why. Gate 6 measured this build's parts on two other images:

  • GGUF on sync-0.34.4-908, whose payload is hash-identical to the release image's;
  • MLX and OCRBench on the candidate, whose MLX payload is the release image's.

So those numbers held for production by argument. The run re-measured on maxusai/ollama:sync-0.34.4-main (bf59c2eab9b8), the image :11497 runs, checked by image ID. It ran in canaries with production's environment, from 00:19 to 06:54 today. 18 suites and 8 OCRBench runs: 0 errors, 0 OOMs, 0 not converged.

What it found:

  • GGUF, 8 models: 0 of 6,908 cells move against the 908 image. Against 0.34.2, the deployed build moves exactly the cells the 908 image moves. The batch, KV type and image decoding match the 908 image's run.
  • MLX, 5 models, two runs: 3.0 % of cells per pair against the candidate, which is the candidate's own rate in gate 6. The two per-model counts past the same-build top are both from run 1, and run 2 does not repeat them. Gate 6 found the same two models reaching past its range.
  • OCRBench, rows 0–200, gemma4:31b, four arms, n = 2: every arm equals 0.34.2 item for item. GGUF q4 recovers the Artistic Text item the unpatched candidate lost, and equals the device-half arm (908).

Also in the section:

  • Why the two v0.34.4 variables cannot move a think-off cell. OLLAMA_KV_CACHE_TYPE reaches only llama-server launches. Every served parser returns no thinking-close string while thinking is off. Two-pass does not defer a think-off request.
  • The manifest digest of all 15 measured models (ADR 0038).
  • What the check does not cover: think-on, OCRBench beyond 200 rows (the Metal host ran all 1000 on its build, fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375), and timing.

How the tables were made (ADR 0012, SPEC H7). Every table is generator output, pasted by script. Cell counts come from cellcmp-0344.py, one call per pair. OCRBench tables come from summarize_extbench.py --repeats --categories and --paired, verbatim; the q4 table's MIXED banner is intended, since three builds are the point. The suite generator's provenance footer reads one host and 0.34.4-dynres-0-gb43ee8e per leg, and all 26 score files carry the stamp.

Checks:

ai-server/mlx-cuda

🤖 Generated with Claude Code

…qual to the 908 image, MLX in its spread, OCRBench equal

Gate 6 measured the build's parts on two other images: GGUF on sync-0.34.4-908, MLX and OCRBench on the candidate.
Those numbers held for production by argument. As 0.34.2's record did, this re-runs the think-off suite and the
OCRBench arms on the image :11497 serves (sync-0.34.4-main, bf59c2eab9b8, 0.34.4-dynres-0-gb43ee8e), in canaries
with production's environment (knob 0, KV f16, two-pass), 2026-09-29 00:19-06:54:
- GGUF, 8 models: 0 of 6,908 cells against the 908 image; against 0.34.2 the same cells the 908 image moves.
- MLX, 5 models, two runs: 3.0 % of cells per pair against the candidate, the candidate's own gate 6 rate; the two
  counts past the same-build top are run 1's and do not repeat.
- OCRBench rows 0-200, four arms, n = 2: every arm equals 0.34.2 item for item; q4 recovers the item the fold lost.
The section records why the two v0.34.4 variables cannot move a think-off cell, the models' manifest digests
(ADR 0038) and what the check does not cover; the status table's deploy row links it. Tables are rendered by
cellcmp-0344.py and summarize_extbench.py and pasted by script.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit 2a32595 into main Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant