Skip to content

unsloth: repin DiffusionGemma and Inkling to their conflict-fixed heads - #230

Merged
danielhanchen merged 1 commit into
masterfrom
repin-24423-25731
Sep 22, 2026
Merged

danielhanchen merged 1 commit into
masterfrom
repin-24423-25731

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

The nightly prebuild has failed at Resolve tag every run since 09-19. Both pins below stopped merging onto the base tag, and one refusal aborts the whole merge, so nothing is built.

What broke

Upstream ggml-org#29042 (landed 09-18) taught the model saver to write the SWA pattern and removed 15 architectures from llama_model_saver_supports_arch(). That is the same block both pins add their own architecture to. Upstream deleting text that a pin edits is not a pure add/add, so additive_merge.py refuses it, exactly as designed:

refused  src/llama-model-saver.cpp: merge base is not empty, so at least one side edited existing text
ggml-org/llama.cpp#24423 (12e0a96) does not merge cleanly onto b11078 + the PRs listed before it

ggml-org#25731 had two further collisions: it and upstream both claimed vocab pre-type id 59, and upstream rewrote the sparse indices index math in the same two lines where the pin added its banded-bias pointer.

What changed here

Both PR branches now carry a merge of current upstream master with those conflicts resolved, so the pins move to the new heads:

pin old new
ggml-org#24423 DiffusionGemma 12e0a96 3dac51d
ggml-org#25731 TML Inkling 946fc11 3870aa4

The resolutions keep upstream's deletion and keep each architecture excluded from the saver for the reason that is still live: add_kv_from_model does not write diffusion.canvas_length or the inkling.* hparams, so a saved model of either cannot be loaded back. The old comments blamed the SWA pattern, which upstream has now fixed. Inkling's vocab pre-type moves to the next free id, which is safe because the id is internal and the GGUF carries tokenizer.ggml.pre as a string.

Verification

  • The full pin set merges onto b11078 with no refusal: DiffusionGemma ggml-org/llama.cpp#24423 clean, Add TML Inkling architecture ggml-org/llama.cpp#25731 additive, everything after unchanged.
  • pin_contract.py reports all 13 pins intact in the merged tree.
  • Both PR heads build with CUDA and pass test-llama-archs: diffusion-gemma NMSE vs CPU 3.66e-07 on B200 and 0.00e+00 on CPU, inkling 1.98e-07 and 0.00e+00.
  • Real models run on the merged tree: diffusiongemma-26B-A4B-it-Q4_K_M at 200.6 tok/s through llama-diffusion-cli, and Inkling-Small-UD-IQ1_S at about 70 tok/s through llama-cli with flash attention on, matching its output with flash attention off.
  • Both upstream PRs report MERGEABLE again.

Both pins have been stale since upstream ggml-org#29042 landed on
09-18. That commit taught the model saver to write the SWA pattern and removed
15 architectures from llama_model_saver_supports_arch(), which is the same
block both pins add their architecture to. Upstream deleting text a pin edits
is not a pure add/add, so additive_merge.py refuses it and the resolve step
fails. Every nightly from 09-19 on has died there.

Both PR branches now carry a merge of current upstream master with the
conflicts resolved, so the pins move to those heads:

  ggml-org#24423 12e0a96 -> 3dac51d
  ggml-org#25731 946fc11 -> 3870aa4

The resolutions keep upstream's deletion and keep each architecture excluded
from the saver for the reason that is still true: add_kv_from_model does not
write diffusion.canvas_length or the inkling.* hparams, so a saved model of
either cannot be loaded back. The old comments blamed the SWA pattern, which
upstream has now fixed.

Verified: the full pin set merges onto b11078 with no refusal (ggml-org#24423 clean,
ggml-org#25731 additive), pin_contract.py reports all 13 pins intact, and both
architectures build and pass test-llama-archs on CUDA and CPU.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-22T12:35:35.511152Z 5cd9a39 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@danielhanchen
danielhanchen merged commit 44953b2 into master Sep 22, 2026
0 of 5 checks passed
@danielhanchen
danielhanchen deleted the repin-24423-25731 branch September 22, 2026 13:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant