The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
The official GGUF conversion of Qwen2.5-0.5B-Instruct, quantised to
Q4_K_M: 491 MB, ready for llama.cpp, Ollama, LM Studio and anything else that speaks GGUF.
Same 494,032,768-parameter model as the bf16 checkpoint, converted by the Qwen team rather than by a
third party, and quantised to the Q4_K_M recipe — 4-bit weights with a mixed-precision scheme that
keeps the attention and feed-forward projections at higher precision than the rest.
Q4_K_M is the usual default quantisation for a reason: about 4.4 bits per weight, a small and
tolerable quality loss, and a file half the size of the bf16 original. Other quantisations in the same
repository run from Q3_K_M at 432 MB up to fp16 at 1.27 GB.
Identical to the upstream instruct model: pre-trained on 18 trillion tokens, then supervised
fine-tuning and preference optimisation. This repository only re-encodes those weights for CPU
inference. It inherits every property and every limitation of the model it is converted from.
Good for: CPU-only machines, laptops without a usable GPU, embedded devices, and bundling inside a
desktop app that cannot ship a Python stack.
Do not rely on it for: the same things the bf16 model cannot do, and expect slightly worse output
again because of the quantisation. Q4_K_M is a good trade, not a free one — quality degrades fastest
on rare tokens, code and non-English text.