Popular repositories Loading
-
sglang-ascend-glm53-serving
sglang-ascend-glm53-serving PublicGLM-5.3 (GlmMoeDsa 256E) production inference on Ascend 910B with SGLang: EP32×DP-attention×NEXTN, cross-node DeepEP patch, W4A8 MoE fast path — 1552 tok/s @1024 concurrent. Agent-readable skill pa…
Python
-
sglang-ascend-glm53-flash-serving
sglang-ascend-glm53-flash-serving PublicGLM-5.3-Flash (glm5_next: hybrid KDA+DSA, 288-expert MoE) production inference on Ascend 910B with SGLang: Python-level port from zero upstream support, msmodelslim fused-MoE W8A8/W4A8 quantization…
Python
-
-
v100-dgx2-qwen38-flash-next-serving
v100-dgx2-qwen38-flash-next-serving PublicQwen3.8-Flash-Next (qwen4exp, 180B/6B active) production serving on 8x V100-DGX2: llama.cpp -sm tensor TP4 (arch blacklist removal + CUDA graph cache cap fix, 12→44 t/s single, 130 t/s @ 8-slot con…
Python
If the problem persists, check the GitHub status page or contact support.