Request: DFlash draft checkpoint for ornith-ai/Ornith-1.5-35B-A3B
Model: ornith-ai/Ornith-1.5-35B-A3B (qwen3_5_moe arch, 35B total / 3B active, hybrid GDN linear attention + full attention)
Why this model
Ornith-1.5-35B-A3B is a popular recent open-weights fine-tune of the Qwen3.5/3.6 35B-A3B base, and it's a strong fit for local agent orchestration on Apple Silicon (oMLX reports active community usage).
Compatibility is already verified
I tested the existing z-lab/Qwen3.6-35B-A3B-DFlash drafter against the Ornith target on oMLX 0.6.3rc1 (DFlash 2 runtime):
- Vocab matches exactly: 248320 = 248320 ✓
- Tokenizer/chat template compatible ✓
- Runtime loads cleanly, no corruption ✓
But cross-model acceptance is poor
Same-machine, same-load A/B (M5 Pro 64GB, oMLX 0.6.3rc1, oQ4e-fp16-mtp target):
| Engine |
Decode speed |
| Native MTP head (bundled with Ornith) |
72-85 tok/s |
| DFlash with Qwen3.6-35B-A3B drafter |
46-49 tok/s |
The distribution mismatch from Ornith's fine-tune costs ~33% — a properly distilled drafter for Ornith itself should flip this the other way (your published numbers show DFlash 2 at +31% acceptance over native MTP).
Ask
A DFlash/DFlash 2 draft checkpoint distilled from Ornith-1.5-35B-A3B itself (like the existing Qwen3.6-35B-A3B one). Happy to contribute validation data / acceptance-length measurements on Apple Silicon once it's out.
Thank you for DFlash — the oMLX integration is excellent.
Request: DFlash draft checkpoint for ornith-ai/Ornith-1.5-35B-A3B
Model: ornith-ai/Ornith-1.5-35B-A3B (qwen3_5_moe arch, 35B total / 3B active, hybrid GDN linear attention + full attention)
Why this model
Ornith-1.5-35B-A3B is a popular recent open-weights fine-tune of the Qwen3.5/3.6 35B-A3B base, and it's a strong fit for local agent orchestration on Apple Silicon (oMLX reports active community usage).
Compatibility is already verified
I tested the existing
z-lab/Qwen3.6-35B-A3B-DFlashdrafter against the Ornith target on oMLX 0.6.3rc1 (DFlash 2 runtime):But cross-model acceptance is poor
Same-machine, same-load A/B (M5 Pro 64GB, oMLX 0.6.3rc1, oQ4e-fp16-mtp target):
The distribution mismatch from Ornith's fine-tune costs ~33% — a properly distilled drafter for Ornith itself should flip this the other way (your published numbers show DFlash 2 at +31% acceptance over native MTP).
Ask
A DFlash/DFlash 2 draft checkpoint distilled from Ornith-1.5-35B-A3B itself (like the existing Qwen3.6-35B-A3B one). Happy to contribute validation data / acceptance-length measurements on Apple Silicon once it's out.
Thank you for DFlash — the oMLX integration is excellent.