docs: v0.34.4 tagged, released and deployed on CUDA — fold record, README, production preflight run - #390
Conversation
…CUDA section - Status: tag and deploy. v0.34.4-dynres is cut on b43ee8e, the #375 merge (ADR 0032), so the release build stamps 0.34.4-dynres-0-gb43ee8e, which cuda-dynres-903's version_pattern already accepts. The tag push started release.yaml, which stopped at its first step as designed (no self-hosted runners). The CUDA deploy stays prepared, not run. - The CUDA think-on section still quoted Metal's interim gemma4:26b count ("five more"); Metal's final count is 4 against 1 (#375). #389 fixed the gfx1151 section's copy; this fixes the other. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
gfx1151: the release image is built and gated, at the maintainer's word. Nothing is deployed.
Deploying is the maintainer's call, and so is the flow: single pass, or
|
maxusai/ollama:sync-0.34.4-main (bf59c2eab9b8) is the gated sync-0.34.4-908 with a Go-only swap to the 0.34.4-dynres-0-gb43ee8e binary, as 0.34.2's release was. Nothing a native stage copies changed since 5584539 but llama/compat/README.md; all 2,696 payload files are hash-identical to -908's. Preflight on a canary from the tag's harness: VERDICT PASS, PASS=21 SKIP=8, every check's status equal to run 3's on -908. The deploy waits on the maintainer's word. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
gfx1151 production is on v0.34.4 with the two-pass flow, at the maintainer's word. It was deployed at 07:39 AEST, 21:39 UTC on 2026-09-27.
For the fold record's "tag and deploy" row: gfx1151 is deployed, with the same two-pass setting that CUDA's prepared deploy uses.
|
…uction preflight run - :11497 runs ollama-0.34.4-dynres-0-gb43ee8e (sync-0.34.4-main) since 2026-09-28 07:37:54, on the maintainer's word: 13 s without service, 55 models either side, OLLAMA_KV_CACHE_TYPE=f16 and OLLAMA_FORMAT_TWO_PASS=1 added to the mirrored environment. Post-deploy preflight on production: VERDICT PASS, PASS=21 SKIP=8, every check's status equal to the canary's; all 9 KV allocations f16; the startup config reads OLLAMA_FORMAT_TWO_PASS:true. The 0.34.2 container is kept for rollback. - The run is committed (runs/preflight-cuda-0344-prod-gb43ee8e.json, with git add -f: runs/ is ignored), as the README's generator comment asked, so the release matrix now regenerates both production rows: release_matrix.py --version 0.34.4-dynres runs/*.json (cuda, and rocm7 from #391). - README: the current fold is v0.34.4-dynres; the Deployed block names the 0.34.4 build on CUDA and gfx1151 with its two variables; the obsolete 0.34.1 gfx1151 paragraph (all three hosts on one commit, patch 906, retired in 0.34.2) is removed, its history is in amd-upgrade-gate.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
CUDA production is on v0.34.4 with the two-pass flow, at the maintainer's word. It was deployed at 07:37:54 AEST, 21:37 UTC on 2026-09-27, two minutes before gfx1151.
|
|
The maintainer merged #390 into
|
Follow-up to #375, which the maintainer merged as
b43ee8e37. It records v0.34.4 on CUDA from tag to production, in the fold record, the README and one committed preflight run.v0.34.4-dynresis cut onb43ee8e37, the fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375 merge, per ADR 0032. So a release build stamps0.34.4-dynres-0-gb43ee8e, whichcuda-dynres-903'sversion_patternaccepts. The tag push startedrelease.yaml, which stopped at its first step ("Require self-hosted runners before building a release"), as designed.maxusai/ollama:sync-0.34.4-main(bf59c2eab9b8) is the gatedsync-0.34.4-908with a Go-only swap to the0.34.4-dynres-0-gb43ee8ebinary.5584539, nothing a native stage copies has changed exceptllama/compat/README.md.-908's.:11497runsollama-0.34.4-dynres-0-gb43ee8e.OLLAMA_KV_CACHE_TYPE=f16andOLLAMA_FORMAT_TWO_PASS=1(open items 6 and 7). The startup config readsOLLAMA_FORMAT_TWO_PASS:true, and all 9 KV allocations are f16.preflight/runs/preflight-cuda-0344-prod-gb43ee8e.json;runs/is ignored, so it was added withgit add -f), as the README's generator comment asked. The release matrix now regenerates both production rows withrelease_matrix.py --version 0.34.4-dynres runs/*.json:cudafrom this run,rocm7from gate: 0.34.4 promoted on gfx1151, with the two-pass flow #391's.v0.34.4-dynres.amd-upgrade-gate.md.Not in this PR: the README's upstream-comparison tables are still written against v0.34.1. The
think+formatrow in particular predates 0.34.4's single pass. That rewrite belongs with the ADR that supersedes ADR 0004 (open item 3).ai-server/mlx-cuda🤖 Generated with Claude Code