The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
The same 494,032,768-parameter model, with weights quantised to 4 bits and stored in the MLX layout so
that Apple's unified-memory GPUs can run it natively. At 278 MB the whole model fits comfortably in the
memory an M-series machine has spare.
Parameters
494,032,768 (before quantisation)
File
model.safetensors, 278,064,920 bytes
Quantisation
4-bit MLX
Context
32,768 tokens, up to 8,192 generated
Licence
Apache-2.0
MLX quantisation is not the same thing as GGUF quantisation even at the same nominal bit width: the
grouping and the runtime differ, and MLX does not have a CPU fallback story as good as llama.cpp's.
Pick MLX if you are on a Mac, GGUF if you are not.
Not trained here — converted. Pre-training, supervised fine-tuning and preference optimisation all
happened upstream for Qwen2.5-0.5B-Instruct; this repository re-encodes the result. Every limitation of
that model applies, plus whatever the 4-bit conversion costs on long-tail tokens.
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Qwen2.5-0.5B-Instruct-4bit")
print(generate(model, tokenizer, prompt="Summarise the idea of a tiny model hub.", max_tokens=128))
Good for: Mac-only local inference, prototypes that need an OpenAI-compatible endpoint without a
network, and anything memory-bound where 278 MB versus 988 MB is the difference between fitting and
not fitting.
Do not rely on it for: non-Apple hardware, or for output quality equal to the bf16 model. 4-bit
quantisation of a 0.5B model is a compression of something already small; the damage is real.