Skip to content

docs: 908 is a fork choice, not an upstream bug — no lost precision, and the loops follow the rounding - #399

Merged
glennneuber merged 1 commit into
mainfrom
docs/908-a-fork-choice
Sep 28, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/908-a-fork-choice

Conversation

@glennneuber

Copy link
Copy Markdown

This records two checks on plain llama.cpp from 2026-09-28. Together they show that ce8caa6e6, the retuned flash-attention tiling behind patch 908, is not an upstream bug. It changes docs only.

Both builds are plain llama.cpp: b11081 (161755f2) as released, and the same tree with only 908 applied. Both were built against CUDA 13.0 for 120-virtual, as production's payload is, with no other fork patch.

  • The loop does not travel. I gave ggml-org's public gemma-4-26B-A4B Q4_0 the exact multi_3img_anchored request ollama's runner sent: greedy, with production's llama-server flags.

    • b11081 finishes in 8,342 tokens with every question right.
    • The revert loops ("Let's re-estimate." ×92).

    With ollama's Q4_K_M it is the other way round.

  • The precision is unchanged. I ran llama.cpp's own test-backend-ops: FLASH_ATTN_EXT against the CPU reference, over the default cases at head sizes 256 and 512, with every case's NMSE printed.

    • Both builds pass all 263 cases.
    • In every head-size × batch group, the paired error ratio released/reverted is 0.96–1.02.

So the retune loses no accuracy. Which case loops follows the rounding and the weights, as #387 found for the KV types. 908 stays: it keeps the numerics production was gated on, which also decode faster on gemma4:26b. Nothing is filed upstream.

The changes:

  • Fold record: a new section, "908 against upstream", with both checks' generator output verbatim. Item 2 now says 908 is a fork choice.
  • README: the 908 row states the accuracy result and that the loop direction depends on the weights. The closing paragraph no longer implies an upstream report is due.
  • Retirement register: the 908 row now retires on a fold's own loop-rate run instead of an upstream retune.

The anchors, links and source-path check are clean, and the name scan of the diff finds nothing.

ai-server/mlx-cuda

🤖 Generated with Claude Code

…and the loops follow the rounding

Two checks on plain llama.cpp (2026-09-28): b11081 (161755f2) as released
and the same tree with only 908, both built against CUDA 13.0 for
120-virtual as production's payload is.
- The loop does not travel. ggml-org's public gemma-4-26B-A4B Q4_0 with
  the exact multi_3img_anchored request ollama's runner sent (greedy,
  production's llama-server flags): b11081 finishes in 8,342 tokens with
  every question right; the revert loops ("Let's re-estimate." x92).
- The precision is unchanged. test-backend-ops FLASH_ATTN_EXT against the
  CPU reference, default cases at head sizes 256/512, every case's NMSE
  printed: 263 cases pass on both, and the paired error ratio
  released/reverted is 0.96-1.02 in every head-size x batch group.

So the retune loses no accuracy; which case loops follows the rounding and
the weights (#387's finding for the KV types). The fold record gains the
evidence (generator output, verbatim) and item 2 says so; the README's 908
row and closing paragraph stop implying an upstream report is due; the
retirement register's 908 row retires on a fold's own loop-rate run
instead of an upstream retune. Nothing is filed upstream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit f36df4e into main Sep 28, 2026
3 checks passed
@glennneuber

Copy link
Copy Markdown
Author

The maintainer merged #399 into main as f36df4ebc. The 908 row and the retirement register now record that the revert is the fork's choice to keep production's gated numerics: ce8caa6e6 loses no precision in test-backend-ops. The row still notes that the revert costs no speed on CUDA and changes no gfx1151 kernel.

amd-server/rocm-gfx1151

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant