Skip to content

add GLM-5.3-Flash (GLM5-Next) support - #27773

Open
timkhronos wants to merge 57 commits into
ggml-org:masterfrom
timkhronos:GLM5.3-Flash
Open

timkhronos wants to merge 57 commits into
ggml-org:masterfrom
timkhronos:GLM5.3-Flash

Conversation

@timkhronos

@timkhronos timkhronos commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Add support for GLM -5.3-flash a 320B hybrid model, supporting both text and vision.

Additional information

Architecture

GLM 5.3 flash mixes 34 KDA linear layers with 11 DSA laters, with mHC and Deepseek style Moe. Most of the parts are already in llama.cpp so I reused whatever I could:

  • KDA layers reuse the Kimi-K3 implementation
  • Attention layers are nope only MLA
  • For mHC I reused the Deepseek V4 implementation. I moved the build_hc helpers from the DSV4 graph into graph_context so both models can share them.
  • Moe and swiglu clamping follow DSV4.

What I implemented new:

  • Here, the DSA indexer scores pools of 4 consecutive token, and always keeps the incomplete tail. I implemented this on top of the existing DSA cache. The indexer cache stores key | gate per token and pooling happens in the graph. No new backend ops have been added.
  • llama_memory_hybrid_dsa: recurrent state + DSA cache, cloned from hybrid ISWA. Rebased onto llama_memory_hybrid_idx instead of the earlier ISWA clone.
  • Vision Tower: the encoder is the same family as glmv4 with per head qk-norm, clamped Swiglu and no post conv norm. It reuses glm4v projector with a swiglu_limit key and an optional image token budget. Added as glm5v as GLM 5.3 Flash requires a different pre processing method than what glm4v uses.
  • Small, precision sensitive tensors (indexer, mHC mixers, KDA gates, MLA low rank paths, roughly 1GB total) are kept unquantized.

Tests

  • Logits match transformers on a small random model across full prefill, small ubatches and single token decode while sparse selection is active, covering both scatter and gather.
  • Vision embeddings match to ~1e-5.
  • The converted model generates coherently and correctly, and vision is working as expected.

Limitations

Quantized GGUFs converted with this PR are available here.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, AI was used in an assistive capacity, and helped figure out and solve several conversion issues, and helped validate the correctness of the implementation.

@github-actions github-actions Bot added model Model specific mtmd Related to multimodal functionality (video/image/audio) conversion labels Aug 26, 2026
Comment thread tools/mtmd/clip.cpp Outdated
Comment thread tools/mtmd/clip-model.h Outdated
Comment thread tools/mtmd/clip-impl.h Outdated
@danielhanchen

danielhanchen commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Hey @timkhronos great work on the PR! A few requests if possible:

  1. You're using glm5-next for general.architecture, whilst model: add GLM-5-Next (GLM-5.3-Flash) #27754 and model : add GLM-5.3-Flash (glm5next) #27752 uses glm5next - Qwen3-Next for eg does qwen3next - I'm unsure what the convention is @ngxson but adopting glm5next might be more generalized? If glm5-next is accepted, a simple first shard rewrite for our uploads should suffice.

  2. The bigger issue is blk.N.indexer.kpool_ape / kpool_gate vs blk.N.indexer_compressor_ape / _gate. deepseek4 mainline already uses blk.N.indexer_compressor_ape and blk.N.indexer_compressor_gate but your PR changes it - the ones we uploaded uses deepseek4's convention. If this PR is accepted, we have to provide a script to rewrite all tensor names or folks have to re-download. If not, can you add aliases so the ones we published works - thanks in advance. Seems like it's more complex than I expected.

Tagging @ngxson for visibility as well.

I re-checked and if (1) + (2) is applied, the quants we uploaded work fine (+ the small shard-1 rewrite) and KLD / PPL are correct under this PR.

Seems like a simple alias isn't possible actually :( It breaks the quants made with this PR

@danielhanchen

Copy link
Copy Markdown
Contributor

Hmmm https://github.com/timkhronos/llama.cpp/pull/9/changes would alias the tensors but it looks a bit problematic hmmm

@Sciguy429

Copy link
Copy Markdown

Throwing up some performance numbers here from the lower end of consumer hardware (128GB DDR5 + 24GB VRAM (4090)).

avar6 has some freshly converted imatrix quants from this PR up as of now if anyone else wants to give them a go: https://huggingface.co/avar6/GLM-5.3-Flash-BF16-gguf

For the IQ3_S, I am getting roughly 300t/s prefill at 256K context and 2048 b/ub size. Generation speed starts off at around 9t/s and drops down considerably by mid window (~128K) to around 6t/s. This seems to track with the 'pooled indexer keys' issue. The model is fully coherent and seems to be working fine. I don't have PPL/KL numbers at the moment as I still need to generate a logit dump.

I have noticed an interesting memory quirk, which I haven't seen before. This is the only model I have ever seen have inconsistent checkpoint sizes. As the context fills the checkpoints grow alarmingly fast in size. At ~90K they are already up to nearly 1.6GB. I don't know if this is an expected behavior for this model arch, or if this is a something which needs to be looked into.

Also, something of note for you @danielhanchen which I found last night while looking over the three PRs for this arch. The vision towers between this PR and yours differ as well. This PR reuses the name clip.vision.projector_type = "glm4v" while you built a new one clip.vision.projector_type = "glm5next". Likely not much of an issue given how easy it is to regenerate mmproj files, but it will need to be delt with as well.

@danielhanchen

Copy link
Copy Markdown
Contributor

Yes I'll re-do the vision! This is fine!

@timkhronos I confirmed timkhronos#9 works fine and does not break your GGUFs. We will however need to do a cheap shard-1 update so that should be fine

@danielhanchen

Copy link
Copy Markdown
Contributor

@timkhronos I saw you changed the tensor naming - but my solution I provided was to allow everyone's quants to work - now your own ones you uploaded don't work haha.

We still need to provide the shard rewrite for the naming (glm5-next) which we're fine with, but now the DeepSeek convention means you yourself have to reupload all shards or do a tensor rename inplace with a script - was this your intention?

@timkhronos

timkhronos commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@danielhanchen Hey!

I ended up going with the the indexer_compressor naming scheme, as it is closer to what's already there, and I was meaning to ask Avar to reconvert anyways, as his ggufs were made when we were missing quantization protection for some crucial tensors so they are not ideal.

Your vision projectors will need reconverting though most likely, and your main model ggufs might be missing the index_share_for_mtp_iteration key as well.

@danielhanchen

Copy link
Copy Markdown
Contributor

@timkhronos Hey! I made some shard rewrites to https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/main/Shard_Rewrite for in preparation!

@am17an am17an left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we add a test for this memory type? or maybe enroll it in test-recurrent-state-rollback once seq_rm works without repooling the entire kv-cache

Comment thread src/llama-memory-hybrid-idx.h Outdated
// The pooled keys persist in the idx cache across batches.
// Sequence edits shift the pool grid and stale the cached values.
bool mem_idx_is_stale() const { return mem_idx_stale; }
void mem_idx_stale_clear () { mem_idx_stale = false; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
void mem_idx_stale_clear () { mem_idx_stale = false; }
void mem_idx_stale_clear() { mem_idx_stale = false; }

Comment thread tools/mtmd/clip.cpp Outdated
Comment on lines 651 to 656
if (hparams.has_swiglu_clamp()) {
cur = ggml_clamp(ctx0, cur, hparams.swiglu_clamp_gate.first, hparams.swiglu_clamp_gate.second);
tmp = ggml_clamp(ctx0, tmp, hparams.swiglu_clamp_up.first, hparams.swiglu_clamp_up.second);
cb(cur, "ffn_gate_clamped", il);
}
cur = ggml_swiglu_split(ctx0, cur, tmp);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we already have ggml_swiglu_clamp, please re-use that

Comment thread src/models/glm5-next.cpp Outdated

ggml_tensor * pooled_new = nullptr;
// Pool only entries completed by this ubatch.
if (n_new > 0) {

@am17an am17an Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is going to have a graph thrash every k_pool tokens, it's better to write a dummy value regardless of n_new

Comment thread src/models/glm5-next.cpp Outdated
pooled_new = ggml_reshape_2d(ctx0, pooled_new, n_embd_indexer, n_new);
cb(pooled_new, "indexer_pool_k_new", il);

if (inp_kpool->cache_safe) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should try to eliminate the concept of cache_safe as it will change graph topology and generally have quite a performance cost relative to the "safe" option where no two seqs share a cell. ATM it seems like the only way to really avoid this to not use the kv-unified cache. cc: @ggerganov

Comment thread src/models/models.h
graph(const llm_graph_params & params) : llm_graph_context(params) {}
graph(const llama_model & model, const llm_graph_params & params);
// manifold-constrained hyper-connections (mHC), shared by deepseek4 and derived model graphs like glm5-next.
template <typename Base = llm_graph_context>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this abstraction maybe too early, as every model right now is creating it's own slightly different version of these ops. It may be better to keep these separate for now and refactor later

Comment thread src/llama-memory-hybrid-idx.cpp Outdated
new llama_kv_cache_context(mem->get_mem_idx(), std::move(sinfos_idx), ubatches)) {}
new llama_kv_cache_context(mem->get_mem_idx(), std::move(sinfos_idx), ubatches)) {
// Sequence edits require a full re-pool.
mem_idx_stale_batch = mem->mem_idx_is_stale();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

right now any edit to the mem_idx is going to cause a full re-pool. Since we're keeping the entire kv cache along with the pool entries, why do we need to do this? We can just re-do the pool from that point? This will allow MTP to roll-back effectively. seq_rm from tail should be free in this setup

Comment thread src/llama-hparams.h Outdated
uint32_t indexer_top_k = 0;
uint32_t indexer_kpool = 0; // k-pool size
bool indexer_kpool_select_tail = true;
bool indexer_index_share_mtp = false; // MTP iterations reuse one indexer selection

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks like this is not used at the moment

Comment thread src/llama-memory-hybrid-idx.cpp Outdated
Comment on lines +815 to +834
// Pools start at the first valid token
for (size_t j = 0; j + kpool <= sq.cells.size(); ) {
const llama_pos p0 = sq.cells[j].first;
if ((p0 - sq.pos_min) % (llama_pos) kpool != 0) {
++j;
continue;
}
bool ok = true;
for (uint32_t k = 1; k < kpool; ++k) {
if (sq.cells[j + k].first != p0 + (llama_pos) k) {
ok = false;
break;
}
}
if (ok) {
sq.pools.push_back((uint32_t) j);
j += kpool;
} else {
++j;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like this machinery exists in case there are holes in the sequence. This usually is only required for context shift iirc, maybe switching it off will simplify this code a lot. Not required in this PR though

@am17an

am17an commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Also test-save-load-state should work with this model


// Fix the pool layout of this ubatch.
if (res && kpool_track()) {
kpool_st = std::make_unique<kpool_state>(kpool_build_state(get_ubatch()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is going to repack the entire pool state for every ubatch, we should fix this

The two clamps around swiglu_split are what ggml_swiglu_clamp already does,
so the clamp bounds collapse back to one value. GLM5V also never called
set_limit_image_tokens(), so --image-max-tokens had no effect.

Assisted-by: Claude Opus 5
(cherry picked from commit 46d18e1)
The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed
by the positions relative to the sequence's first one, so the layout now lives
on the memory and a ubatch only appends to it.

A sequence edit no longer stales every pooled key either, only the ones at or
after the edited position, which makes a tail seq_rm free. The pooling subgraph
is built unconditionally so the graph shape no longer changes every kpool
tokens, and the pool axis is folded into rows before soft_max, which otherwise
exceeds the CUDA gridDim.y limit past n_kv 262144.

Assisted-by: Claude Opus 5
(cherry picked from commit 5d1c40b)
The conv state and the delta net state were only written to the live row, so a
rollback restored whatever the checkpoint rows happened to hold. Take the same
route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv
groups. That also drops a state view that assumed contiguous rows.

Enroll the arch in test-recurrent-state-rollback, which catches this under its
garbage-filled cache pass.

Assisted-by: Claude Opus 5
(cherry picked from commit 5ace37e)
…flicts

Merge ggml-org/llama.cpp master (a97cce8) into the GLM-5.3-Flash
(GLM5-Next) branch of PR ggml-org#27773 so the branch becomes mergeable again
(GitHub reported mergeable=false, mergeable_state=dirty).

Exactly one file conflicted; all other files merged automatically:

* src/llama-graph.cpp (build_moe_ffn swiglu-clamp path): both sides
  added a new architecture to the same OR-chain that selects the fused
  ggml_swiglu_clamp path. master added LLM_ARCH_MAPLE; this branch added
  LLM_ARCH_GLM5_NEXT. Resolved by keeping BOTH arches in the condition:
      if (arch == LLM_ARCH_MAPLE || arch == LLM_ARCH_DEEPSEEK4 ||
          arch == LLM_ARCH_GLM5_NEXT ||
          (arch == LLM_ARCH_DFLASH && hparams.dsv4_hc_mult > 0) ||
          arch == LLM_ARCH_HY_V4)
  No other change to the file; the branch's separate GLM5_NEXT addition
  in build_ffn (swiglu_clamp_shexp) auto-merged untouched.

Assisted-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@pwilkin

pwilkin commented Sep 27, 2026

Copy link
Copy Markdown
Member

@am17an @ngxson PTAL when you have a moment.

Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence.

Assisted-by: Codex
Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error.

Assisted-by: Codex
@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

CI is failing

@pwilkin

pwilkin commented Sep 28, 2026

Copy link
Copy Markdown
Member

Yeah I'm on it.

@pwilkin

pwilkin commented Sep 28, 2026

Copy link
Copy Markdown
Member

No idea what's with the Windows checks but I fixed the other ones.

@am17an

am17an commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Probably you need to rebase as well. Not sure if this failure is caused by this PR https://github.com/ggml-org/llama.cpp/actions/runs/36382726074/job/108808733541?pr=27773

…teardown

Two defects in the cross-ubatch k-pool layout added by the k-pool commit:

1. Wrong results. An edited sequence only rebuilt its pool layout when its cell
   count changed, so if the first ubatch after an edit added back exactly as many
   cells as were removed, the stale position-to-cell list survived. With a unified
   cache and more than one sequence, where another sequence takes the freed cells,
   the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN).
   Rebuild whenever the sequence is stale, not only on a size mismatch.

2. Slowdown. "shared" mode was assumed to end only with an edit that forces a
   rebuild, but sharing also ends when the other sequence is removed. The survivor
   kept shared = true, pinning cache_safe off and re-pooling every pool on every
   ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in
   copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence
   so one rebuild re-derives sharing and cache_safe returns to 1.

Assisted-by: Claude Opus 5

@am17an am17an left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can support -sm tensor and MTP in follow-up PRs but this looks good to me. Ran a couple of tests locally.

@ggerganov

Copy link
Copy Markdown
Member

Could you rebase/merge master?

@ggerganov ggerganov self-assigned this Sep 28, 2026
build_attn_mha split the batch into streams with a stream stride of
q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2],
when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part
and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the
stride was then n_head times too large and every stream s >= 1 read another
head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was
wrong for every stream past the first. Unified KV and decode were unaffected
(n_stream == 1, and decode takes the gather path). Other MLA models concat rope
so q is contiguous and the computed value is unchanged for them.

Compute the stride from the token dimension, which is identical for a
contiguous q.

Assisted-by: Claude Opus 5
The shared-cell teardown added to seq_rm (stale every sequence when the layout
has shared cells, so a survivor does not keep shared = true and pin cache_safe
off) was missing from the other paths that can free shared cells: state_read
and state_drop staled only the one sequence. Apply the same re-derivation there
and correct the comment that claimed sharing ends only via an edit or seq_rm.

Assisted-by: Claude Opus 5
The hc_ name filter was listed twice in the GLM5_NEXT protection block.

Assisted-by: Claude Opus 5
Bring the PR current with ggml-org master (f00a64c). The only conflict was in
tests/CMakeLists.txt: master's ggml-org#29426 refactored test-recurrent-state-rollback
to run across every generated model via --models, which already covers glm5-next
(it is generated by test-llama-archs and is recurrent, so not skipped, and the
runner fails if any model fails). The PR's separate per-arch glm5-next
invocation was therefore redundant and dropped; glm5-next rollback is still
exercised by the all-archs run.
@pwilkin

pwilkin commented Sep 28, 2026

Copy link
Copy Markdown
Member

Done :)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion model Model specific mtmd Related to multimodal functionality (video/image/audio) testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.