Conversation
Replace the inline DFlash2-only special case in load_draft with a table keyed on architectures[0], adding MuseGlimmerAssistantModel (Meta's v1 DFlash head: encoder.fc/encoder.output_norm_enc key remap, top-level config fields, forced bidirectional in-block attention, default vocab 202048). DFlashConfig.num_target_layers is now optional since the v1 checkpoint never carries it. bind() also pulls output_multiplier/final_logit_softcapping from the target's args when the draft config is left at its structural default (1.0 / None), so the v1 head's logits land on Glimmer's scale. Qwen3 targets have no such attrs and are unaffected.
shoemoney
added a commit
to shoemoney/dflash
that referenced
this pull request
Sep 8, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds Muse-Glimmer-30B to the MLX backend. Meta's
meta-models/Muse-Glimmer-30B-assistant(DFlash v1) loads as published, no weight conversion, alongsidez-lab/Muse-Glimmer-30B-DFlash2, against mlx-lm'smuse_glimmertarget (mlx-lm git main; the PyPI release predates Glimmer).Why
The README lists Muse-Glimmer-30B for Transformers only. On Apple Silicon there was no way to run either head. Meta's checkpoint is Qwen3-shaped (same layout
DFlashDraftModelalready implements; vLLM'sMuseGlimmerAssistantConfigsubclassesQwen3Configfor the same reason), so it only needs what vLLM also patches in:encoder.fc.->fc.,encoder.output_norm_enc.->hidden_norm.vocab_sizeis_causalFalse(bidirectional in-block attention; the loader would otherwise default the sliding layers to causal)How
load_draftnow dispatches through_DRAFT_ADAPTERS, a dict keyed onconfig.architectures[0]mapping to (model class, config normalizer, weight remap). The existingDFlash2DraftModelhandling moves into its entry unchanged;MuseGlimmerAssistantModelis the new entry. Unknown architectures raise with the name.DFlashConfig.num_target_layersbecomes optional (only used whentarget_layer_idsis absent).bind()inherits the target'soutput_multiplier/final_logit_softcappingwhen the draft config leaves them at defaults, so v1 draft probs sit on the target's scale for rejection sampling. Qwen targets have neither attribute and are unaffected.dflash/bench_mlx.py: with/without-draft benchmark writing one JSONL row per prompt and block size (tok/s, accepted tokens per step, peak memory, full token list and its SHA-256), with a--baselineoption.tests/test_muse_glimmer_adapter.py: weight-free. Fixtures are the realconfig.jsonand safetensors headers of both heads; asserts the remapped key set equalsDFlashDraftModel/DFlash2DraftModelparameter keys, plus config fields and the unknown-architecture error. 7 tests, ~1 s.Generate loop, attention, rejection sampler and cache trimming are untouched.
Measured
M3 Ultra (103 GB), mlx 0.32.2, mlx-lm 0.32.0 (git main),
mlx-community/Muse-Glimmer-30B-4bit, greedy,Reasoning strength: low, 6 prompts x 256 new tokens, one request at a time.mlx_lm.stream_generate)Block 5 is the crossover on a quantized target, matching the README's existing
block_size <= 5note. On an abliterated 4-bit Glimmer, acceptance for both heads is within 0.1 of the stock model.One note for reviewers: greedy token equality with plain decoding does not hold on this target in bf16. Replaying the baseline's own tokens through the target alone, single-step vs one batched pass, flips the argmax at near-tie positions (top-2 margin 0.0 to 0.4 under the tanh cap of 20). So the bench reports SHAs but the correctness evidence is by-construction verification plus healthy acceptance, not a token-exact match.
Under sampling (T=1, top-p 0.95, top-k 64) the v1 head accepts only 1.16 tokens/step with the current independent per-position proposal; that is a separate, behavior-changing fix, in #169.