The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
A dense decoder-only transformer with grouped-query attention, released as part of the Qwen3 family
(0.6B through 235B). Unusually for a model this small, it carries the family's hybrid reasoning
behaviour: a thinking mode that emits a reasoning trace before the answer, and a non-thinking mode
that answers directly. The mode is switched by the chat template, not by a separate checkpoint.
Parameters
751,632,384 (includes the 151,936-token vocabulary embeddings)
File
model.safetensors, 1,503,300,328 bytes (bf16)
Context
32,768 tokens natively; extensible with YaRN
Languages
100+
Licence
Apache-2.0
The parameter count is large next to the model's apparent size because Qwen3 uses a big multilingual
vocabulary. Most of that 751M sits in the embedding table; the transformer body is far smaller, which
is why 0.6B feels like a smaller model than the numbers suggest.
Qwen3 was pre-trained on roughly 36 trillion tokens across 119 languages and dialects, then
post-trained in a multi-stage pipeline that includes a strong-to-weak distillation pass to bring
reasoning ability down into the small sizes. This checkpoint is the post-trained model; its base is
Qwen/Qwen3-0.6B-Base.
Good for: on-device chat, small reasoning tasks, multilingual responses, and anything that benefits
from a visible reasoning trace you can inspect.
Do not rely on it for: hard mathematics or multi-step planning — leaving thinking mode on makes it
slower without making it right. The 0.6B size still fails at tasks a 4B model handles, and the
thinking traces are not a reliable account of how the answer was produced.