docs: 908 is a fork choice, not an upstream bug — no lost precision, and the loops follow the rounding - #399
Merged
Merged
Conversation
…and the loops follow the rounding
Two checks on plain llama.cpp (2026-09-28): b11081 (161755f2) as released
and the same tree with only 908, both built against CUDA 13.0 for
120-virtual as production's payload is.
- The loop does not travel. ggml-org's public gemma-4-26B-A4B Q4_0 with
the exact multi_3img_anchored request ollama's runner sent (greedy,
production's llama-server flags): b11081 finishes in 8,342 tokens with
every question right; the revert loops ("Let's re-estimate." x92).
- The precision is unchanged. test-backend-ops FLASH_ATTN_EXT against the
CPU reference, default cases at head sizes 256/512, every case's NMSE
printed: 263 cases pass on both, and the paired error ratio
released/reverted is 0.96-1.02 in every head-size x batch group.
So the retune loses no accuracy; which case loops follows the rounding and
the weights (#387's finding for the KV types). The fold record gains the
evidence (generator output, verbatim) and item 2 says so; the README's 908
row and closing paragraph stop implying an upstream report is due; the
retirement register's 908 row retires on a fold's own loop-rate run
instead of an upstream retune. Nothing is filed upstream.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
|
The maintainer merged #399 into
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This records two checks on plain llama.cpp from 2026-09-28. Together they show that
ce8caa6e6, the retuned flash-attention tiling behind patch 908, is not an upstream bug. It changes docs only.Both builds are plain llama.cpp:
b11081(161755f2) as released, and the same tree with only 908 applied. Both were built against CUDA 13.0 for120-virtual, as production's payload is, with no other fork patch.The loop does not travel. I gave ggml-org's public gemma-4-26B-A4B Q4_0 the exact
multi_3img_anchoredrequest ollama's runner sent: greedy, with production's llama-server flags.b11081finishes in 8,342 tokens with every question right.With ollama's Q4_K_M it is the other way round.
The precision is unchanged. I ran llama.cpp's own
test-backend-ops:FLASH_ATTN_EXTagainst the CPU reference, over the default cases at head sizes 256 and 512, with every case's NMSE printed.So the retune loses no accuracy. Which case loops follows the rounding and the weights, as #387 found for the KV types. 908 stays: it keeps the numerics production was gated on, which also decode faster on gemma4:26b. Nothing is filed upstream.
The changes:
The anchors, links and source-path check are clean, and the name scan of the diff finds nothing.
ai-server/mlx-cuda🤖 Generated with Claude Code