docs(fold): v0.34.4 verified on CUDA's deployed build itself - #411
Merged
Merged
Conversation
…qual to the 908 image, MLX in its spread, OCRBench equal Gate 6 measured the build's parts on two other images: GGUF on sync-0.34.4-908, MLX and OCRBench on the candidate. Those numbers held for production by argument. As 0.34.2's record did, this re-runs the think-off suite and the OCRBench arms on the image :11497 serves (sync-0.34.4-main, bf59c2eab9b8, 0.34.4-dynres-0-gb43ee8e), in canaries with production's environment (knob 0, KV f16, two-pass), 2026-09-29 00:19-06:54: - GGUF, 8 models: 0 of 6,908 cells against the 908 image; against 0.34.2 the same cells the 908 image moves. - MLX, 5 models, two runs: 3.0 % of cells per pair against the candidate, the candidate's own gate 6 rate; the two counts past the same-build top are run 1's and do not repeat. - OCRBench rows 0-200, four arms, n = 2: every arm equals 0.34.2 item for item; q4 recovers the item the fold lost. The section records why the two v0.34.4 variables cannot move a think-off cell, the models' manifest digests (ADR 0038) and what the check does not cover; the status table's deploy row links it. Tables are rendered by cellcmp-0344.py and summarize_extbench.py and pasted by script. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds a section to the v0.34.4 fold record: "Deployed, and verified on the build itself: CUDA (2026-09-29)". The status table's "tag and deploy" row gets one sentence linking to it. It follows 0.34.2's record, which re-measured the shipped image after its deploy. Docs only, one file.
Why. Gate 6 measured this build's parts on two other images:
sync-0.34.4-908, whose payload is hash-identical to the release image's;So those numbers held for production by argument. The run re-measured on
maxusai/ollama:sync-0.34.4-main(bf59c2eab9b8), the image:11497runs, checked by image ID. It ran in canaries with production's environment, from 00:19 to 06:54 today. 18 suites and 8 OCRBench runs: 0 errors, 0 OOMs, 0 not converged.What it found:
Also in the section:
OLLAMA_KV_CACHE_TYPEreaches only llama-server launches. Every served parser returns no thinking-close string while thinking is off. Two-pass does not defer a think-off request.How the tables were made (ADR 0012, SPEC H7). Every table is generator output, pasted by script. Cell counts come from
cellcmp-0344.py, one call per pair. OCRBench tables come fromsummarize_extbench.py --repeats --categoriesand--paired, verbatim; the q4 table's MIXED banner is intended, since three builds are the point. The suite generator's provenance footer reads one host and0.34.4-dynres-0-gb43ee8eper leg, and all 26 score files carry the stamp.Checks:
check_source_paths.py --changed-since origin/mainis clean.ai-server/mlx-cuda🤖 Generated with Claude Code