The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
The smallest instruct model in Alibaba's Qwen2.5 line: 494M parameters, under a gigabyte in bf16, and
surprisingly competent at structured output for its size.
A decoder-only transformer with grouped-query attention and SwiGLU activations, post-trained to follow
instructions and hold a chat. Qwen2.5-0.5B is the base checkpoint; this repository is the instruction-tuned
build.
Parameters
494,032,768
File
model.safetensors, 988,097,824 bytes (bf16)
Context
32,768 tokens, up to 8,192 generated
Languages
English plus roughly two dozen others
Licence
Apache-2.0
Qwen2.5 brings a noticeable jump over Qwen2 in knowledge, coding and mathematics, and in following a
system prompt. It is also much better at emitting valid JSON, which is the usual reason to reach for a
model this small.
The Qwen2.5 family was pre-trained on a corpus of 18 trillion tokens with improved filtering and
deduplication over Qwen2, then post-trained with supervised fine-tuning and preference optimisation.
Only the 0.5B size is relevant here; the family runs up to 72B.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-0.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [
{"role": "system", "content": "You extract fields and answer with JSON only."},
{"role": "user", "content": "Extract the city and the year: 'I moved to Lisbon in 2019.'"},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Good for: chat, tagging, extraction, small classification-ish tasks, JSON-shaped output, and as a
local backend for a tool that must not call an API.
Do not rely on it for: deep reasoning, long documents, reliable arithmetic, or anything where a wrong
answer is expensive. Treat the "two dozen languages" claim as "it can respond in them", not "it is
fluent in them". It inherits the biases of its web-scale training data.