tinymodels
Qwen

Qwen2.5-0.5B-Instruct-GGUF

The same 0.5B instruct model as a 491 MB Q4_K_M GGUF, ready for llama.cpp and Ollama.

text-generation gguf 494M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q4_k_m.gguf. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 64-token reply to a fixed promptgreedy decoding, llama.cpp; one warm-up run then one timed run 255.4 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 795 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

Qwen2.5-0.5B-Instruct (GGUF, Q4_K_M)

The official GGUF conversion of Qwen2.5-0.5B-Instruct, quantised to Q4_K_M: 491 MB, ready for llama.cpp, Ollama, LM Studio and anything else that speaks GGUF.

What it is

Same 494,032,768-parameter model as the bf16 checkpoint, converted by the Qwen team rather than by a third party, and quantised to the Q4_K_M recipe — 4-bit weights with a mixed-precision scheme that keeps the attention and feed-forward projections at higher precision than the rest.

Parameters494,032,768 (before quantisation)
Fileqwen2.5-0.5b-instruct-q4_k_m.gguf, 491,400,032 bytes
QuantisationQ4_K_M
Context32,768 tokens, up to 8,192 generated
LicenceApache-2.0

Q4_K_M is the usual default quantisation for a reason: about 4.4 bits per weight, a small and tolerable quality loss, and a file half the size of the bf16 original. Other quantisations in the same repository run from Q3_K_M at 432 MB up to fp16 at 1.27 GB.

How it was trained

Identical to the upstream instruct model: pre-trained on 18 trillion tokens, then supervised fine-tuning and preference optimisation. This repository only re-encodes those weights for CPU inference. It inherits every property and every limitation of the model it is converted from.

Reference: Qwen2.5 Technical Report, https://arxiv.org/abs/2407.10671.

How to run it

With llama.cpp:

llama-cli -hf Qwen/Qwen2.5-0.5B-Instruct-GGUF:Q4_K_M -p "Name three uses for a tiny language model." -n 128

With Ollama, a one-line Modelfile:

FROM https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q4_k_m.gguf

then ollama create qwen-tiny -f Modelfile && ollama run qwen-tiny.

With Python bindings:

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="Qwen/Qwen2.5-0.5B-Instruct-GGUF",
    filename="qwen2.5-0.5b-instruct-q4_k_m.gguf",
    n_ctx=4096,
)
print(llm.create_chat_completion(
    messages=[{"role": "user", "content": "Explain KV caching in two sentences."}],
))

Intended use and limits

Good for: CPU-only machines, laptops without a usable GPU, embedded devices, and bundling inside a desktop app that cannot ship a Python stack.

Do not rely on it for: the same things the bf16 model cannot do, and expect slightly worse output again because of the quantisation. Q4_K_M is a good trade, not a free one — quality degrades fastest on rare tokens, code and non-English text.

Upstream