gate: 0.34.4 promoted on gfx1151, with the two-pass flow - #391
Merged
Merged
Conversation
ollama-rocm moved from 0.34.3-dynres-0-g650f8fd to 0.34.4-dynres-0-gb43ee8e at 07:39 on 2026-09-28, on the maintainer's word, with OLLAMA_FORMAT_TWO_PASS=1 as on CUDA. The payload moved from b10969 to b11081, with compat 908. - Status: image, version, build (908 added), payload, a think+format flow row, and the retained 0.34.3 container (f16) for rollback. - The decision: rollback, why the fold's gfx1151 evidence covers the promoted image (native payload byte-identical to the 908 image; payload structure equal; Go source equal but for a test file), the q8_0-vs-f16 caveat for think-on, preflight PASS=20 SKIP=12 on the release image and on the promoted container, and the multi_3img_anchored loop that 0.34.3 already had. - Clause 3 is satisfied under (b), and under (c) by transitivity only: a direct direct-io on/off A/B on b11081 was not run. - The compose .env still names 0.32.1; the decision says not to redeploy through compose until it names this image. - The two preflight run records. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
glennneuber
added a commit
that referenced
this pull request
Sep 27, 2026
…uction preflight run - :11497 runs ollama-0.34.4-dynres-0-gb43ee8e (sync-0.34.4-main) since 2026-09-28 07:37:54, on the maintainer's word: 13 s without service, 55 models either side, OLLAMA_KV_CACHE_TYPE=f16 and OLLAMA_FORMAT_TWO_PASS=1 added to the mirrored environment. Post-deploy preflight on production: VERDICT PASS, PASS=21 SKIP=8, every check's status equal to the canary's; all 9 KV allocations f16; the startup config reads OLLAMA_FORMAT_TWO_PASS:true. The 0.34.2 container is kept for rollback. - The run is committed (runs/preflight-cuda-0344-prod-gb43ee8e.json, with git add -f: runs/ is ignored), as the README's generator comment asked, so the release matrix now regenerates both production rows: release_matrix.py --version 0.34.4-dynres runs/*.json (cuda, and rocm7 from #391). - README: the current fold is v0.34.4-dynres; the Deployed block names the 0.34.4 build on CUDA and gfx1151 with its two variables; the obsolete 0.34.1 gfx1151 paragraph (all three hosts on one commit, patch 906, retired in 0.34.2) is removed, its history is in amd-upgrade-gate.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This records the gfx1151 production promotion in
docs/maxusai/amd-upgrade-gate.md, as #386 recorded the KV change.What happened
ollama-rocm(:11434) moved from0.34.3-dynres-0-g650f8fdto0.34.4-dynres-0-gb43ee8eat 07:39 on 2026-09-28, on the maintainer's word. It was down for about one second.OLLAMA_FORMAT_TWO_PASS=1), as on CUDA (fold: upstream v0.34.4 — llama.cpp b11081, MLX 59d600b5, XGrammar 0.2.7 #375, item 7).ollama-rocm-0.34.3-650f8fda(f16,--restart no). The decision lists the rollback commands.Why the fold's gfx1151 evidence covers the promoted image
llama-serverand the four native libraries.payload_diff.shfinds 1863 = 1863 entries and 96/96 gfx1151 kernels.rocm7-0-34-4-dynres,--quality): PASS=20 SKIP=12 on the release image in a bench container, and again on the promoted container itself. On production, every model load shows--cache-type-k f16 --cache-type-v f16 --flash-attn on. Both run records are included.Stated plainly in the decision
--direct-io) is satisfied under (b). Under (c) it holds only by transitivity: b11081 with direct-io on equals b10969 with it on, and b10969's on/off A/B was equal on 2026-09-21. A direct on/off A/B on b11081 was not run.q8_0, which was production's KV type at the time. Think-on cells are not expected to reproduce exactly under f16 (kv: production runs an f16 KV cache on every platform, and a cross-host loop test #387).multi_3img_anchoredunder f16 and greedy decoding. 0.34.3 had the same loop, byte for byte..envstill names 0.32.1. Do not redeploy through compose until its image line is updated. The compose file itself now sets two-pass (MaxusAI/ollama-deployments3354a42).Verification
check_source_paths.py --changed-since origin/mainis clean.amd-server/rocm-gfx1151🤖 Generated with Claude Code