tinymodels
Qwen

Qwen3-0.6B

The 0.6B Qwen3 dense model: on-device chat with a switchable thinking mode.

text-generation safetensors 752M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/Qwen/Qwen3-0.6B/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 64-token reply to a fixed promptgreedy decoding, mlx-lm on the GPU, publisher bf16 weights; one warm-up run then one timed run 469.6 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 1,697 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
MMLU-Redux, thinking modeQwen3 Technical Report, Qwen3-0.6B row of the thinking-mode table 55.6% not stated arxiv.org
MATH-500, thinking modesame table 77.6% not stated arxiv.org
IFEval (strict prompt), thinking modesame table 59.2% not stated arxiv.org

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

Qwen3-0.6B

The 0.6B dense model from the Qwen3 generation: 751M parameters, a toggle between "thinking" and direct answers, and a 32K context window.

What it is

A dense decoder-only transformer with grouped-query attention, released as part of the Qwen3 family (0.6B through 235B). Unusually for a model this small, it carries the family's hybrid reasoning behaviour: a thinking mode that emits a reasoning trace before the answer, and a non-thinking mode that answers directly. The mode is switched by the chat template, not by a separate checkpoint.

Parameters751,632,384 (includes the 151,936-token vocabulary embeddings)
Filemodel.safetensors, 1,503,300,328 bytes (bf16)
Context32,768 tokens natively; extensible with YaRN
Languages100+
LicenceApache-2.0

The parameter count is large next to the model's apparent size because Qwen3 uses a big multilingual vocabulary. Most of that 751M sits in the embedding table; the transformer body is far smaller, which is why 0.6B feels like a smaller model than the numbers suggest.

How it was trained

Qwen3 was pre-trained on roughly 36 trillion tokens across 119 languages and dialects, then post-trained in a multi-stage pipeline that includes a strong-to-weak distillation pass to bring reasoning ability down into the small sizes. This checkpoint is the post-trained model; its base is Qwen/Qwen3-0.6B-Base.

Reference: Qwen3 Technical Report, https://arxiv.org/abs/2505.09388.

How to run it

from transformers import AutoModelForCausalLM, AutoTokenizer

name = "Qwen/Qwen3-0.6B"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype="auto", device_map="auto")

messages = [{"role": "user", "content": "How many r's are in 'strawberry'?"}]
# thinking mode on (default):
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
                                     enable_thinking=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Set enable_thinking=False for direct answers, which is what you want in a latency-sensitive loop.

Intended use and limits

Good for: on-device chat, small reasoning tasks, multilingual responses, and anything that benefits from a visible reasoning trace you can inspect.

Do not rely on it for: hard mathematics or multi-step planning — leaving thinking mode on makes it slower without making it right. The 0.6B size still fails at tasks a 4B model handles, and the thinking traces are not a reliable account of how the answer was produced.

Upstream