Add SigLIP 2 base-256 Core ML conversion - #109
Merged
Merged
Conversation
Converts google/siglip2-base-patch16-256 to two fp16 Core ML packages (image and text encoders, L2-normalized outputs) so labels can be given as text at runtime. ImageNet-1k zero-shot on all 50,000 test images, CPU + Neural Engine: 76.76% vs 76.79% for fp32 PyTorch under the same single-prompt protocol, 99.32% identical top-1. Image encoder runs 100% on the ANE at 5.2 ms (3.5 ms GPU) vs 19.1 ms PyTorch MPS on an M5 Pro.
Core ML fp16 on CPU + Neural Engine 94.77% vs PyTorch fp32 94.74% on all 3,669 test images, 99.89% identical top-1, 5.6 ms per image.
Same photos, prompts, and scoring as FluidUse ImageSortCheck. M5 Pro: Core ML 36.3 s (202 photos/s, 94.26%, 262 MB peak) vs transformers fp32 MPS batch 32 102.2 s (72 photos/s, 94.11%, 4.24 GB peak).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Converts
google/siglip2-base-patch16-256to fp16 Core ML image and text encoders (L2-normalized 768-d outputs) for zero-shot image classification. Published: FluidInference/siglip2-base-patch16-256-coreml.Scripts:
convert-coreml.py,compare-models.py,score-pets.py,bench-pytorch-pets.py; reports inreports/. Swift runtime and demo: FluidInference/FluidUse#18.🤖 Generated with Claude Code