tinymodels
Qwen

Qwen2.5-0.5B-Instruct

Half a billion parameters of general chat and instruction following, under 1 GB in bf16.

text-generation safetensors 494M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 64-token reply to a fixed promptgreedy decoding, mlx-lm on the GPU, publisher bf16 weights; one warm-up run then one timed run 266.8 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 1,451 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
Generation throughput, transformersBF16, batch size 1, 2048 tokens generated from a 1-token prompt 47.4 tok/s NVIDIA A100 80GB qwen.readthedocs.io
Generation throughput, vLLMBF16, same conditions 311.55 tok/s NVIDIA A100 80GB qwen.readthedocs.io
GPU memory, transformers BF16at a 1-token input length 0.97 GB NVIDIA A100 80GB qwen.readthedocs.io

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

Qwen2.5-0.5B-Instruct

The smallest instruct model in Alibaba's Qwen2.5 line: 494M parameters, under a gigabyte in bf16, and surprisingly competent at structured output for its size.

What it is

A decoder-only transformer with grouped-query attention and SwiGLU activations, post-trained to follow instructions and hold a chat. Qwen2.5-0.5B is the base checkpoint; this repository is the instruction-tuned build.

Parameters494,032,768
Filemodel.safetensors, 988,097,824 bytes (bf16)
Context32,768 tokens, up to 8,192 generated
LanguagesEnglish plus roughly two dozen others
LicenceApache-2.0

Qwen2.5 brings a noticeable jump over Qwen2 in knowledge, coding and mathematics, and in following a system prompt. It is also much better at emitting valid JSON, which is the usual reason to reach for a model this small.

How it was trained

The Qwen2.5 family was pre-trained on a corpus of 18 trillion tokens with improved filtering and deduplication over Qwen2, then post-trained with supervised fine-tuning and preference optimisation. Only the 0.5B size is relevant here; the family runs up to 72B.

Reference: Qwen2.5 Technical Report, https://arxiv.org/abs/2407.10671.

How to run it

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen2.5-0.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

messages = [
    {"role": "system", "content": "You extract fields and answer with JSON only."},
    {"role": "user", "content": "Extract the city and the year: 'I moved to Lisbon in 2019.'"},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Intended use and limits

Good for: chat, tagging, extraction, small classification-ish tasks, JSON-shaped output, and as a local backend for a tool that must not call an API.

Do not rely on it for: deep reasoning, long documents, reliable arithmetic, or anything where a wrong answer is expensive. Treat the "two dozen languages" claim as "it can respond in them", not "it is fluent in them". It inherits the biases of its web-scale training data.

Upstream