Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
100 changes: 100 additions & 0 deletions models/emb/siglip2/coreml/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
# SigLIP 2 Core ML

Core ML conversion of Google's [SigLIP 2](https://arxiv.org/abs/2502.14786)
image and text encoders, starting with
[`google/siglip2-base-patch16-256`](https://huggingface.co/google/siglip2-base-patch16-256)
(375M parameters: 93M image encoder, 282M text encoder of which 197M is the
256k-token embedding table). Apache-2.0, same as upstream.

Both towers are exported, so labels can be written as text at runtime
(zero-shot classification) instead of being fixed at conversion time.
Published packages: [FluidInference/siglip2-base-patch16-256-coreml](https://huggingface.co/FluidInference/siglip2-base-patch16-256-coreml).

## Packages

| Package | Input | Output |
| --- | --- | --- |
| `siglip2-base-patch16-256-image-fp16.mlpackage` | `pixel_values` float32 `[1, 3, 256, 256]` | `image_embeds` float32 `[1, 768]`, L2-normalized |
| `siglip2-base-patch16-256-text-fp16.mlpackage` | `input_ids` int32 `[1, 64]` | `text_embeds` float32 `[1, 768]`, L2-normalized |

`config.json` records the preprocessing and the scoring constants.

- **Image:** resize to 256×256 (bilinear), scale to [0, 1], then
`(x - 0.5) / 0.5` per channel.
- **Text:** lowercase, Gemma tokenizer, pad to 64 tokens (`padding="max_length"`).
SigLIP was trained without an attention mask, so padding must match.
- **Score:** `sigmoid(logit_scale * cos + logit_bias)` gives an independent
probability per label; `argmax(cos)` gives a single choice. Text embeddings
depend only on the labels, so compute them once and reuse them for every
image.

## Usage

```bash
uv sync
uv run python convert-coreml.py # build/siglip2-base-patch16-256/
uv run python compare-models.py --limit 1000 # Core ML vs PyTorch on ImageNet-1k
uv run python score-pets.py # Core ML vs PyTorch on Oxford-IIIT Pets
uv run python bench-pytorch-pets.py --batch 32 # PyTorch timing baseline for FluidUse ImageSortCheck
```

`compare-models.py` downloads the ImageNet-1k test split from
[`clip-benchmark/wds_imagenet1k`](https://huggingface.co/datasets/clip-benchmark/wds_imagenet1k)
(about 6.4 GB) and scores zero-shot classification with the prompt
`this is a photo of {class}.`, lowercased.

## Validation

All numbers from an Apple M5 Pro (macOS 27).

**ImageNet-1k zero-shot, all 50,000 test images**, fp16 packages on
CPU + Neural Engine, against the fp32 PyTorch model under the same protocol
([report](reports/imagenet1k-zeroshot-base-256-fp16-ane.json)):

| | PyTorch fp32 | Core ML fp16 |
| --- | ---: | ---: |
| Top-1 | 76.79% | 76.76% |
| Same prediction as PyTorch | — | 99.32% |
| Image embedding cosine, mean / min | — | 0.99990 / 0.975 |
| Text embedding cosine (1,000 class prompts), mean / min | — | 0.99998 / 0.9989 |

Google reports 79.1% for this checkpoint with its own class names and prompts.
The single-prompt protocol here scores lower for both backends; the table
compares Core ML with PyTorch, not with Google's number.

**Latency per call** (batch 1, median):

| Backend | Image encoder | Text encoder |
| --- | ---: | ---: |
| Core ML, CPU + Neural Engine (100% of ops on ANE for image) | 5.2 ms | 1.3 ms |
| Core ML, CPU + GPU | 3.5 ms | 3.0 ms |
| Core ML, CPU only | 17.4 ms | 4.1 ms |
| PyTorch fp32, MPS | 19.1 ms | — |
| PyTorch fp32, CPU | 91.4 ms | — |

**Oxford-IIIT Pets** (`score-pets.py`, 3,669 test photos, 37 breeds, prompt
`a photo of a {breed}, a type of pet.`): Core ML 94.77%, PyTorch 94.74%, 99.89% identical top-1
([report](reports/oxford-pets-zeroshot-base-256-fp16-ane.json)).

**End to end against PyTorch** (`bench-pytorch-pets.py` vs the FluidUse Swift `ImageSortCheck`, 7,349 Pets
photos, decode + resize + encoder + scoring, after model load;
[report](reports/pets-7349-coreml-vs-pytorch.json)):

| | Core ML (Swift, 4 in flight) | transformers fp32, MPS, batch 32 |
| --- | ---: | ---: |
| Time | 36.3 s | 102.2 s |
| Photos per second | 202 | 72 |
| Peak memory | 262 MB | 4.24 GB |
| Model on disk | 715 MB | 1.50 GB |
| Accuracy | 94.26% | 94.11% |

## Notes

- The converter warns about fp16 overflow while folding constants; the
comparison above shows it does not affect outputs.
- `coreml-cli` cannot build a compute plan for the text package (likely the
256k-row embedding gather), but the package loads and runs on every compute
unit.
- NaFlex (variable-resolution) checkpoints are not supported: they need dynamic
shapes, and the SigLIP 2 authors report the fixed-resolution checkpoints are
better on dense tasks.
64 changes: 64 additions & 0 deletions models/emb/siglip2/coreml/bench-pytorch-pets.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
"""PyTorch (transformers) baseline for the Swift ImageSortCheck: same photos, prompts, and scoring."""

import argparse
import json
import os
import time
from pathlib import Path

import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor

BREEDS = [
"abyssinian", "american bulldog", "american pit bull terrier", "basset hound", "beagle", "bengal", "birman",
"bombay", "boxer", "british shorthair", "chihuahua", "egyptian mau", "english cocker spaniel", "english setter",
"german shorthaired", "great pyrenees", "havanese", "japanese chin", "keeshond", "leonberger", "maine coon",
"miniature pinscher", "newfoundland", "persian", "pomeranian", "pug", "ragdoll", "russian blue", "saint bernard",
"samoyed", "scottish terrier", "shiba inu", "siamese", "sphynx", "staffordshire bull terrier", "wheaten terrier",
"yorkshire terrier",
] # fmt: skip


def main():
parser = argparse.ArgumentParser()
parser.add_argument("--device", default="mps")
parser.add_argument("--batch", type=int, default=32)
parser.add_argument("--dtype", default="float32")
args = parser.parse_args()
cache = Path(os.path.expanduser("~/Library/Caches/FluidUse/image-sort/oxford-pets-test"))
labels = {int(k): v for k, v in json.loads((cache / "manifest.json").read_text()).items()}
ids = sorted(labels)
dtype = getattr(torch, args.dtype)
model = AutoModel.from_pretrained("google/siglip2-base-patch16-256", dtype=dtype).eval().to(args.device)
processor = AutoProcessor.from_pretrained("google/siglip2-base-patch16-256")
prompts = [f"a photo of a {b}, a type of pet." for b in BREEDS]
with torch.no_grad():
tokens = processor(text=prompts, padding="max_length", max_length=64, return_tensors="pt").input_ids
text = model.text_model(input_ids=tokens.to(args.device)).pooler_output
text = torch.nn.functional.normalize(text.float(), dim=-1)
# warm up
warm = processor(images=[Image.open(cache / f"{ids[0]}.jpg").convert("RGB")], return_tensors="pt")
model.vision_model(pixel_values=warm.pixel_values.to(args.device, dtype))
if args.device == "mps":
torch.mps.synchronize()
start = time.perf_counter()
correct = 0
for offset in range(0, len(ids), args.batch):
chunk = ids[offset : offset + args.batch]
images = [Image.open(cache / f"{i}.jpg").convert("RGB") for i in chunk]
pixels = processor(images=images, return_tensors="pt").pixel_values.to(args.device, dtype)
emb = torch.nn.functional.normalize(model.vision_model(pixel_values=pixels).pooler_output.float(), dim=-1)
pred = (emb @ text.T).argmax(-1).cpu().tolist()
correct += sum(int(p == labels[i]) for p, i in zip(pred, chunk))
if args.device == "mps":
torch.mps.synchronize()
seconds = time.perf_counter() - start
n = len(ids)
print(json.dumps({"device": args.device, "batch": args.batch, "dtype": args.dtype, "photos": n,
"seconds": round(seconds, 2), "photos_per_s": round(n / seconds, 1),
"ms_per_photo": round(1000 * seconds / n, 2), "accuracy": round(correct / n, 4)}))


if __name__ == "__main__":
main()
129 changes: 129 additions & 0 deletions models/emb/siglip2/coreml/compare-models.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
"""Compare Core ML SigLIP2 encoders with PyTorch on ImageNet-1k zero-shot classification."""

import argparse
import io
import json
import tarfile
import time
from pathlib import Path

import coremltools as ct
import numpy as np
import torch
from huggingface_hub import snapshot_download
from PIL import Image
from transformers import AutoModel, AutoProcessor

DATASET = "clip-benchmark/wds_imagenet1k"
COMPUTE_UNITS = {
"all": ct.ComputeUnit.ALL,
"cpu_and_ne": ct.ComputeUnit.CPU_AND_NE,
"cpu_and_gpu": ct.ComputeUnit.CPU_AND_GPU,
"cpu": ct.ComputeUnit.CPU_ONLY,
}


def iter_samples(root: Path, limit: int):
shards = sorted((root / "test").glob("*.tar"), key=lambda p: int(p.stem))
count = 0
for shard in shards:
with tarfile.open(shard) as tar:
pending = {}
for member in tar:
key, _, ext = member.name.rpartition(".")
pending.setdefault(key, {})[ext] = tar.extractfile(member).read()
sample = pending[key]
image_ext = next((e for e in ("jpg", "jpeg", "png", "webp") if e in sample), None)
if image_ext and "cls" in sample:
del pending[key]
image = Image.open(io.BytesIO(sample[image_ext])).convert("RGB")
yield image, int(sample["cls"].decode().strip())
count += 1
if count >= limit:
return


def text_prompts(root: Path, template: str):
classnames = (root / "classnames.txt").read_text().splitlines()
return [template.format(c).lower() for c in classnames if c]


def main():
parser = argparse.ArgumentParser()
parser.add_argument("--build-dir", type=Path, default=Path("build/siglip2-base-patch16-256"))
parser.add_argument("--precision", default="fp16")
parser.add_argument("--limit", type=int, default=1000)
parser.add_argument("--compute-units", default="all", choices=COMPUTE_UNITS)
parser.add_argument("--template", default="this is a photo of {}.")
parser.add_argument("--report", type=Path)
args = parser.parse_args()

config = json.loads((args.build_dir / "config.json").read_text())
name = config["model_id"].split("/")[-1]
units = COMPUTE_UNITS[args.compute_units]
image_model = ct.models.MLModel(
str(args.build_dir / f"{name}-image-{args.precision}.mlpackage"), compute_units=units
)
text_model = ct.models.MLModel(
str(args.build_dir / f"{name}-text-{args.precision}.mlpackage"), compute_units=units
)
model = AutoModel.from_pretrained(config["model_id"], dtype=torch.float32).eval()
processor = AutoProcessor.from_pretrained(config["model_id"])
root = Path(snapshot_download(DATASET, repo_type="dataset", allow_patterns=["test/*", "*.txt"]))

prompts = text_prompts(root, args.template)
tokens = processor(
text=prompts, padding="max_length", max_length=config["text_length"], return_tensors="pt"
).input_ids
with torch.no_grad():
text_ref = model.text_model(input_ids=tokens).pooler_output
text_ref = torch.nn.functional.normalize(text_ref, dim=-1).numpy()
text_cml = np.concatenate(
[
text_model.predict({"input_ids": row[None].numpy().astype(np.int32)})["text_embeds"]
for row in tokens
]
)
text_cos = (text_ref * text_cml).sum(-1)

image_cos, agree, correct_ref, correct_cml, latencies = [], 0, 0, 0, []
for image, label in iter_samples(root, args.limit):
pixels = processor(images=image, return_tensors="pt").pixel_values
with torch.no_grad():
ref = model.vision_model(pixel_values=pixels).pooler_output
ref = torch.nn.functional.normalize(ref, dim=-1)
ref = ref.numpy()[0]
start = time.perf_counter()
cml = image_model.predict({"pixel_values": pixels.numpy()})["image_embeds"][0]
latencies.append((time.perf_counter() - start) * 1000)
image_cos.append(float(ref @ cml))
pred_ref = int(np.argmax(text_ref @ ref))
pred_cml = int(np.argmax(text_cml @ cml))
agree += pred_ref == pred_cml
correct_ref += pred_ref == label
correct_cml += pred_cml == label

n = len(image_cos)
report = {
"model_id": config["model_id"],
"precision": args.precision,
"compute_units": args.compute_units,
"template": args.template,
"images": n,
"text_cosine_min": float(text_cos.min()),
"text_cosine_mean": float(text_cos.mean()),
"image_cosine_min": float(np.min(image_cos)),
"image_cosine_mean": float(np.mean(image_cos)),
"top1_agreement": agree / n,
"top1_pytorch": correct_ref / n,
"top1_coreml": correct_cml / n,
"image_latency_ms_median": float(np.median(latencies[5:] or latencies)),
}
print(json.dumps(report, indent=2))
if args.report:
args.report.parent.mkdir(parents=True, exist_ok=True)
args.report.write_text(json.dumps(report, indent=2) + "\n")


if __name__ == "__main__":
main()
Loading