Skip to content

gate: 0.34.4 promoted on gfx1151, with the two-pass flow - #391

Merged
glennneuber merged 1 commit into
mainfrom
gate/0344-promoted
Sep 27, 2026
Merged

glennneuber merged 1 commit into
mainfrom
gate/0344-promoted

Conversation

@glennneuber

Copy link
Copy Markdown

This records the gfx1151 production promotion in docs/maxusai/amd-upgrade-gate.md, as #386 recorded the KV change.

What happened

ollama-rocm (:11434) moved from 0.34.3-dynres-0-g650f8fd to 0.34.4-dynres-0-gb43ee8e at 07:39 on 2026-09-28, on the maintainer's word. It was down for about one second.

Why the fold's gfx1151 evidence covers the promoted image

  • Native payload: byte-identical to the 908 image that passed the fold's gfx1151 gates. The same sha256 for llama-server and the four native libraries. payload_diff.sh finds 1863 = 1863 entries and 96/96 gfx1151 kernels.
  • Go source: the same as the 908 build, but for one test file.
  • Preflight (rocm7-0-34-4-dynres, --quality): PASS=20 SKIP=12 on the release image in a bench container, and again on the promoted container itself. On production, every model load shows --cache-type-k f16 --cache-type-v f16 --flash-attn on. Both run records are included.

Stated plainly in the decision

  • Clause 3 (--direct-io) is satisfied under (b). Under (c) it holds only by transitivity: b11081 with direct-io on equals b10969 with it on, and b10969's on/off A/B was equal on 2026-09-21. A direct on/off A/B on b11081 was not run.
  • The gate campaigns ran under q8_0, which was production's KV type at the time. Think-on cells are not expected to reproduce exactly under f16 (kv: production runs an f16 KV cache on every platform, and a cross-host loop test #387).
  • gemma4:26b loops on multi_3img_anchored under f16 and greedy decoding. 0.34.3 had the same loop, byte for byte.
  • The compose .env still names 0.32.1. Do not redeploy through compose until its image line is updated. The compose file itself now sets two-pass (MaxusAI/ollama-deployments 3354a42).

Verification

  • check_source_paths.py --changed-since origin/main is clean.
  • The run records contain no host paths or names.

amd-server/rocm-gfx1151

🤖 Generated with Claude Code

ollama-rocm moved from 0.34.3-dynres-0-g650f8fd to 0.34.4-dynres-0-gb43ee8e at
07:39 on 2026-09-28, on the maintainer's word, with OLLAMA_FORMAT_TWO_PASS=1 as
on CUDA. The payload moved from b10969 to b11081, with compat 908.

- Status: image, version, build (908 added), payload, a think+format flow row,
  and the retained 0.34.3 container (f16) for rollback.
- The decision: rollback, why the fold's gfx1151 evidence covers the promoted
  image (native payload byte-identical to the 908 image; payload structure
  equal; Go source equal but for a test file), the q8_0-vs-f16 caveat for
  think-on, preflight PASS=20 SKIP=12 on the release image and on the promoted
  container, and the multi_3img_anchored loop that 0.34.3 already had.
- Clause 3 is satisfied under (b), and under (c) by transitivity only: a direct
  direct-io on/off A/B on b11081 was not run.
- The compose .env still names 0.32.1; the decision says not to redeploy
  through compose until it names this image.
- The two preflight run records.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit a504fcd into main Sep 27, 2026
3 checks passed
glennneuber added a commit that referenced this pull request Sep 27, 2026
…uction preflight run

- :11497 runs ollama-0.34.4-dynres-0-gb43ee8e (sync-0.34.4-main) since
  2026-09-28 07:37:54, on the maintainer's word: 13 s without service,
  55 models either side, OLLAMA_KV_CACHE_TYPE=f16 and
  OLLAMA_FORMAT_TWO_PASS=1 added to the mirrored environment. Post-deploy
  preflight on production: VERDICT PASS, PASS=21 SKIP=8, every check's
  status equal to the canary's; all 9 KV allocations f16; the startup
  config reads OLLAMA_FORMAT_TWO_PASS:true. The 0.34.2 container is kept
  for rollback.
- The run is committed (runs/preflight-cuda-0344-prod-gb43ee8e.json, with
  git add -f: runs/ is ignored), as the README's generator comment asked,
  so the release matrix now regenerates both production rows:
  release_matrix.py --version 0.34.4-dynres runs/*.json (cuda, and rocm7
  from #391).
- README: the current fold is v0.34.4-dynres; the Deployed block names
  the 0.34.4 build on CUDA and gfx1151 with its two variables; the
  obsolete 0.34.1 gfx1151 paragraph (all three hosts on one commit, patch
  906, retired in 0.34.2) is removed, its history is in
  amd-upgrade-gate.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant