Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
1361ba7
feat(pytorch): support GLM-5.3 Flash with shared LMDeploy components
qescccczmr Sep 14, 2026
e2fc086
merge: adapt GLM-5.3 Flash to current main interfaces
qescccczmr Sep 14, 2026
fca7e9f
refactor(pytorch): use default Triton MoE for GLM-5.3 Flash
qescccczmr Sep 17, 2026
d3dbefe
refactor(pytorch): remove unused EP1 DeepGEMM MoE extensions
qescccczmr Sep 17, 2026
6b183f6
merge: resolve main conflicts and adapt GLM indexer prefix
qescccczmr Sep 17, 2026
3bc41ce
merge: sync GLM-5.3 Flash PR with latest main
qescccczmr Sep 21, 2026
6218f59
feat(pytorch): support GLM-5.3 Flash MTP with shared components
qescccczmr Sep 21, 2026
3fcabb7
fix(pytorch): correct GLM-5.3 MTP verification and acceptance
qescccczmr Sep 22, 2026
41563c7
refactor(pytorch): reuse causal convolution cache for GLM MTP
qescccczmr Sep 22, 2026
9642332
fix(pytorch): honor MoE reduction options across backends
qescccczmr Sep 22, 2026
96a10e6
fix(pytorch): address GLM backend and runtime policy review
qescccczmr Sep 23, 2026
a17e90b
style(pytorch): satisfy GLM pre-commit hooks
qescccczmr Sep 23, 2026
0f33060
refactor(pytorch): reuse native LayerNorm for GLM
qescccczmr Sep 23, 2026
17dec3e
refactor: reuse LayerNorm in GLM vision merger
qescccczmr Sep 23, 2026
7b5a819
refactor: reuse common FP32 rotary operator for GLM vision
qescccczmr Sep 23, 2026
c1ae34b
perf: preserve strided KDA decode inputs
qescccczmr Sep 23, 2026
7fa48eb
perf: fuse KDA gates while preserving sigmoid precision
qescccczmr Sep 23, 2026
9d8c5f1
revert: drop DeepGEMM masked GEMM alias compatibility from GLM PR
qescccczmr Sep 23, 2026
7e0b9de
perf: batch GLM KPool prefill cache updates on device
qescccczmr Sep 23, 2026
e542908
perf: batch GLM KPool prefill scoring and selection
qescccczmr Sep 23, 2026
a2800a0
perf: store GLM NoPE latent cache without padding
qescccczmr Sep 23, 2026
868782a
perf: read original HC states in shared pre-reduce
qescccczmr Sep 23, 2026
5988cd5
perf: use compact FP8 MoE scheduling for sparse routes
qescccczmr Sep 23, 2026
82e451c
revert: remove PR-specific test changes
qescccczmr Sep 24, 2026
18d0188
refactor: make MoE weighted reduction use FP32 accumulation
qescccczmr Sep 24, 2026
1a577df
refactor: reuse FlashMLA sparse attention for GLM-5.3
qescccczmr Sep 24, 2026
8068cd4
revert: remove GLM-specific video sampling
Sep 24, 2026
00f8917
revert: remove model media I/O defaults merging
qescccczmr Sep 24, 2026
3f21b23
fix(pytorch): enable GLM DP and expert parallelism
qescccczmr Sep 24, 2026
38779f1
perf(pytorch): fuse GLM KPool verification and reuse shared operators
qescccczmr Sep 28, 2026
e200190
fix(pytorch): accumulate DeepEP local expert reduction in FP32
qescccczmr Sep 28, 2026
67dd161
Merge upstream main into GLM-5.3-Flash support
qescccczmr Sep 28, 2026
8c167eb
fix(pytorch): make DeepEP FP32 local accumulation opt-in
qescccczmr Sep 28, 2026
08d7a74
refactor(pytorch): remove unused sparse MLA kernel and padding hook
qescccczmr Sep 28, 2026
18a8c91
fix(pytorch): preserve MoE reduction defaults with explicit GLM FP32 …
qescccczmr Sep 28, 2026
200a559
fix(pytorch): bound and reserve GLM KPool prefill score memory
qescccczmr Sep 28, 2026
fc843a3
perf(pytorch): skip invalid KPool compression and broadcast prefill g…
qescccczmr Sep 28, 2026
ba4787d
perf: fuse sparse MLA prefill index remapping and padding
qescccczmr Sep 28, 2026
1930759
perf: vectorize mHC input conversion for long prefills
qescccczmr Sep 29, 2026
932285b
style: wrap MoE reduction docstring for lint
qescccczmr Sep 29, 2026
ec2ec44
perf(pytorch): reduce GLM decode projection and metadata overhead
qescccczmr Sep 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions lmdeploy/archs.py
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,7 @@ def check_vl_llm(backend: str, config: dict) -> bool:
'Gemma3ForConditionalGeneration', 'Llama4ForConditionalGeneration', 'InternVLForConditionalGeneration',
'InternS1ForConditionalGeneration', 'InternS1ProForConditionalGeneration',
'InternS1_1_ForConditionalGeneration', 'Glm4vForConditionalGeneration',
'Glm5NextForConditionalGeneration',
'InternS2MobiusForConditionalGeneration', 'InternS2MobiusForCausalLM',
'InternS2PreviewForConditionalGeneration', 'InternS2PreviewForCausalLM',
'KimiK25ForConditionalGeneration', 'Kimi_K25ForConditionalGeneration',
Expand Down
5 changes: 4 additions & 1 deletion lmdeploy/hf_configs/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,10 @@
@lru_cache
def register_config(model_type: str):
"""Register an LMDeploy-owned Transformers config when available."""
if model_type == 'kimi_k2':
if model_type == 'glm5_next':
from .configuration_glm5_next import Glm5NextConfig
AutoConfig.register(Glm5NextConfig.model_type, Glm5NextConfig)
elif model_type == 'kimi_k2':
# Standalone Kimi EAGLE checkpoints do not provide an auto_map.
from .configuration_kimi_k2 import KimiK2Config
AutoConfig.register(KimiK2Config.model_type, KimiK2Config)
Expand Down
Loading
Loading