tinymodels
mlx-community

Qwen2.5-0.5B-Instruct-4bit

4-bit MLX conversion of Qwen2.5-0.5B-Instruct, 278 MB, tuned for Apple silicon.

text-generation mlx 494M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/mlx-community/Qwen2.5-0.5B-Instruct-4bit/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 64-token reply to a fixed promptgreedy decoding, mlx-lm on the GPU, publisher bf16 weights; one warm-up run then one timed run 241.4 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 786 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

Qwen2.5-0.5B-Instruct 4-bit (MLX)

A 4-bit MLX conversion of Qwen2.5-0.5B-Instruct: 278 MB, quantised for Apple silicon, runnable through mlx-lm.

What it is

The same 494,032,768-parameter model, with weights quantised to 4 bits and stored in the MLX layout so that Apple's unified-memory GPUs can run it natively. At 278 MB the whole model fits comfortably in the memory an M-series machine has spare.

Parameters494,032,768 (before quantisation)
Filemodel.safetensors, 278,064,920 bytes
Quantisation4-bit MLX
Context32,768 tokens, up to 8,192 generated
LicenceApache-2.0

MLX quantisation is not the same thing as GGUF quantisation even at the same nominal bit width: the grouping and the runtime differ, and MLX does not have a CPU fallback story as good as llama.cpp's. Pick MLX if you are on a Mac, GGUF if you are not.

How it was trained

Not trained here — converted. Pre-training, supervised fine-tuning and preference optimisation all happened upstream for Qwen2.5-0.5B-Instruct; this repository re-encodes the result. Every limitation of that model applies, plus whatever the 4-bit conversion costs on long-tail tokens.

Reference: Qwen2.5 Technical Report, https://arxiv.org/abs/2407.10671.

How to run it

pip install mlx-lm
mlx_lm.generate --model mlx-community/Qwen2.5-0.5B-Instruct-4bit \
  --prompt "Write a haiku about gradients." --max-tokens 128

Or a local server that speaks the OpenAI protocol, which is the usual reason to reach for this:

mlx_lm.server --model mlx-community/Qwen2.5-0.5B-Instruct-4bit --port 8080
from mlx_lm import load, generate

model, tokenizer = load("mlx-community/Qwen2.5-0.5B-Instruct-4bit")
print(generate(model, tokenizer, prompt="Summarise the idea of a tiny model hub.", max_tokens=128))

Intended use and limits

Good for: Mac-only local inference, prototypes that need an OpenAI-compatible endpoint without a network, and anything memory-bound where 278 MB versus 988 MB is the difference between fitting and not fitting.

Do not rely on it for: non-Apple hardware, or for output quality equal to the bf16 model. 4-bit quantisation of a 0.5B model is a compression of something already small; the damage is real.

Upstream