Skip to content

mlx: load Meta's Muse-Glimmer-30B-assistant DFlash head via a draft-adapter registry - #168

Open
shoemoney wants to merge 2 commits into
z-lab:mainfrom
shoemoney:pr/muse-glimmer-mlx
Open

shoemoney wants to merge 2 commits into
z-lab:mainfrom
shoemoney:pr/muse-glimmer-mlx

Conversation

@shoemoney

@shoemoney shoemoney commented Sep 8, 2026 •

Copy link
Copy Markdown

What

Adds Muse-Glimmer-30B to the MLX backend. Meta's meta-models/Muse-Glimmer-30B-assistant (DFlash v1) loads as published, no weight conversion, alongside z-lab/Muse-Glimmer-30B-DFlash2, against mlx-lm's muse_glimmer target (mlx-lm git main; the PyPI release predates Glimmer).

Why

The README lists Muse-Glimmer-30B for Transformers only. On Apple Silicon there was no way to run either head. Meta's checkpoint is Qwen3-shaped (same layout DFlashDraftModel already implements; vLLM's MuseGlimmerAssistantConfig subclasses Qwen3Config for the same reason), so it only needs what vLLM also patches in:

value
weight remap encoder.fc. -> fc., encoder.output_norm_enc. -> hidden_norm.
vocab_size 202048 (absent from the checkpoint; Qwen3's default would misindex the mask token)
is_causal False (bidirectional in-block attention; the loader would otherwise default the sliding layers to causal)

How

  • load_draft now dispatches through _DRAFT_ADAPTERS, a dict keyed on config.architectures[0] mapping to (model class, config normalizer, weight remap). The existing DFlash2DraftModel handling moves into its entry unchanged; MuseGlimmerAssistantModel is the new entry. Unknown architectures raise with the name. DFlashConfig.num_target_layers becomes optional (only used when target_layer_ids is absent).
  • bind() inherits the target's output_multiplier / final_logit_softcapping when the draft config leaves them at defaults, so v1 draft probs sit on the target's scale for rejection sampling. Qwen targets have neither attribute and are unaffected.
  • dflash/bench_mlx.py: with/without-draft benchmark writing one JSONL row per prompt and block size (tok/s, accepted tokens per step, peak memory, full token list and its SHA-256), with a --baseline option.
  • tests/test_muse_glimmer_adapter.py: weight-free. Fixtures are the real config.json and safetensors headers of both heads; asserts the remapped key set equals DFlashDraftModel / DFlash2DraftModel parameter keys, plus config fields and the unknown-architecture error. 7 tests, ~1 s.
  • README: Muse-Glimmer-30B added to the MLX section with an example and the numbers below.

Generate loop, attention, rejection sampler and cache trimming are untouched.

Measured

M3 Ultra (103 GB), mlx 0.32.2, mlx-lm 0.32.0 (git main), mlx-community/Muse-Glimmer-30B-4bit, greedy, Reasoning strength: low, 6 prompts x 256 new tokens, one request at a time.

draft block tok/s speedup accepted / step
none (plain mlx_lm.stream_generate) - 40.7 1.00x -
Meta assistant, bf16 5 45.8 1.12x 2.99
Meta assistant, 4-bit 5 47.6 1.17x 2.93
Meta assistant, 4-bit 8 40.8 1.00x 3.63
Meta assistant, 4-bit 16 32.6 0.80x 4.68
DFlash 2, 4-bit 5 50.8 1.25x 3.36
DFlash 2, 4-bit 8 45.8 1.13x 4.19
DFlash 2, 4-bit 16 37.0 0.91x 5.47

Block 5 is the crossover on a quantized target, matching the README's existing block_size <= 5 note. On an abliterated 4-bit Glimmer, acceptance for both heads is within 0.1 of the stock model.

One note for reviewers: greedy token equality with plain decoding does not hold on this target in bf16. Replaying the baseline's own tokens through the target alone, single-step vs one batched pass, flips the argmax at near-tie positions (top-2 margin 0.0 to 0.4 under the tanh cap of 20). So the bench reports SHAs but the correctness evidence is by-construction verification plus healthy acceptance, not a token-exact match.

Under sampling (T=1, top-p 0.95, top-k 64) the v1 head accepts only 1.16 tokens/step with the current independent per-position proposal; that is a separate, behavior-changing fix, in #169.

Replace the inline DFlash2-only special case in load_draft with a table
keyed on architectures[0], adding MuseGlimmerAssistantModel (Meta's v1
DFlash head: encoder.fc/encoder.output_norm_enc key remap, top-level
config fields, forced bidirectional in-block attention, default vocab
202048). DFlashConfig.num_target_layers is now optional since the v1
checkpoint never carries it.

bind() also pulls output_multiplier/final_logit_softcapping from the
target's args when the draft config is left at its structural default
(1.0 / None), so the v1 head's logits land on Glimmer's scale. Qwen3
targets have no such attrs and are unaffected.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant