Skip to content

dasLLAMA vulkan: a MoE that fits the card serves on the whole-model resident driver - the routed block on the device, the shared expert's K-quant planes, the decode GEMV lane split; Qwen3.6-35B-A3B whole at 1.05x / 1.50x of llama.cpp, Qwen3-30B-A3B 1.00x / 1.15x, the Qwen1.5-MoE twin 1.03x / 0.96x - #3988

Merged
borisbat merged 22 commits into
masterfrom
bbatkin/qwen-moe-resident
Sep 10, 2026

the review-md discovery test's dasllama fixture carries the surface t…

918b958
Select commit
Loading
Failed to load commit list.
Sign in for the full log view