Skip to content

Request: DFlash draft checkpoint for ornith-ai/Ornith-1.5-35B-A3B #161

Description

@nolanwang1996-cmd

Request: DFlash draft checkpoint for ornith-ai/Ornith-1.5-35B-A3B

Model: ornith-ai/Ornith-1.5-35B-A3B (qwen3_5_moe arch, 35B total / 3B active, hybrid GDN linear attention + full attention)

Why this model

Ornith-1.5-35B-A3B is a popular recent open-weights fine-tune of the Qwen3.5/3.6 35B-A3B base, and it's a strong fit for local agent orchestration on Apple Silicon (oMLX reports active community usage).

Compatibility is already verified

I tested the existing z-lab/Qwen3.6-35B-A3B-DFlash drafter against the Ornith target on oMLX 0.6.3rc1 (DFlash 2 runtime):

  • Vocab matches exactly: 248320 = 248320 ✓
  • Tokenizer/chat template compatible ✓
  • Runtime loads cleanly, no corruption ✓

But cross-model acceptance is poor

Same-machine, same-load A/B (M5 Pro 64GB, oMLX 0.6.3rc1, oQ4e-fp16-mtp target):

Engine Decode speed
Native MTP head (bundled with Ornith) 72-85 tok/s
DFlash with Qwen3.6-35B-A3B drafter 46-49 tok/s

The distribution mismatch from Ornith's fine-tune costs ~33% — a properly distilled drafter for Ornith itself should flip this the other way (your published numbers show DFlash 2 at +31% acceptance over native MTP).

Ask

A DFlash/DFlash 2 draft checkpoint distilled from Ornith-1.5-35B-A3B itself (like the existing Qwen3.6-35B-A3B one). Happy to contribute validation data / acceptance-length measurements on Apple Silicon once it's out.

Thank you for DFlash — the oMLX integration is excellent.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions