Every catalogue model, one fixed input, one machine
The complete measured table for the 2026-09-20 run: one input per task type, one warm-up run then one timed run each, on Apple M4 Pro, 48 GB unified memory, macOS 15.6. This page is the citation for that run — the raw JSON is the machine-readable form of exactly this table.
| model | task | input | ms | peak rss | note |
|---|---|---|---|---|---|
| all-MiniLM-L6-v2 | sentence-similarity | 1 sentence · transformers | 5.3 | 395 MB | one sentence embedded (one forward pass), transformers, float32 |
| distilbert-base-uncased-finetuned-sst-2-english | text-classification | 1 sentence · transformers | 13.1 | 525 MB | one sentence classified (one forward pass), transformers, float32 |
| resnet-50 | image-classification | 1 photo (640x480) · transformers | 21.7 | 505 MB | one forward pass, transformers, float32 |
| SmolLM2-135M-Instruct | text-generation | 64-token reply · mlx-lm | 110.2 | 658 MB | greedy decoding, mlx-lm on the GPU, publisher bf16 weights |
| videomae-base-finetuned-kinetics | video-classification | 1 clip (16 frames) · transformers | 214.2 | 976 MB | one forward pass, transformers, float32, synthetic clip |
| Qwen2.5-0.5B-Instruct-4bit | text-generation | 64-token reply · mlx-lm | 241.4 | 786 MB | greedy decoding, mlx-lm on the GPU, publisher bf16 weights |
| detr-resnet-50 | object-detection | 1 photo (640x480) · transformers | 251.1 | 1373 MB | one forward pass, transformers, float32 |
| Qwen2.5-0.5B-Instruct-GGUF | text-generation | 64-token reply · llama.cpp | 255.4 | 795 MB | greedy decoding, llama.cpp |
| Qwen2.5-0.5B-Instruct | text-generation | 64-token reply · mlx-lm | 266.8 | 1451 MB | greedy decoding, mlx-lm on the GPU, publisher bf16 weights |
| Qwen3-0.6B | text-generation | 64-token reply · mlx-lm | 469.6 | 1697 MB | greedy decoding, mlx-lm on the GPU, publisher bf16 weights |
| Florence-2-base-ft | image-text-to-text | 1 photo + prompt · transformers | 2124.4 | 2109 MB | greedy decoding, transformers, float32 |
| SmolVLM2-256M-Video-Instruct | video-text-to-text | 1 clip + question · transformers | 2841.4 | 3171 MB | greedy decoding, transformers, float32, synthetic clip |
What this is, and what it is not
This is Apple M4 Pro, 48 GB unified memory, macOS 15.6, one laptop, one run per model. It is a latency figure for one fixed input — a 64-token reply, a sentence, a photo, a clip — chosen to be the same kind of work for every model of that task type. It is not a benchmark suite, not an accuracy score, and not a speed ranking across task types: a video model and a chat model are not doing the same work.
Treat the numbers as order-of-magnitude honest: factors of two are noise. One timed run after one warm-up cannot separate two models that are within a couple of times of each other; it can only tell you that a 110 ms reply and a 2,800 ms clip differ in kind. When a figure is quoted anywhere on this hub, it quotes this run and this method.
Re-measuring is one command, bun run measure, which always re-runs every model together — never one model alone, because a number from a different day, runtime or thermal state is not comparable with this table. Each run lands as a new dated page and a new raw JSON; the dated permalinks never change what they point at.
Method text straight from the harness: One fixed input per task type, one warm-up run then one timed run, on the hub's own machine. The harness is scripts/measure.py; the input shapes are defined there and named in each figure.