The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
A 135M-parameter instruction-tuned language model from Hugging Face's SmolLM2 family, built for
on-device chat where a laptop or a browser tab has to do the work.
SmolLM2 is a compact model family at 135M, 360M and 1.7B parameters. This is the smallest instruct
variant: a Llama-style decoder-only transformer with 134,515,008 parameters, stored here as a
269 MB safetensors file in bf16.
Parameters
134,515,008
File
model.safetensors, 269,060,552 bytes
Architecture
Llama-style decoder-only transformer
Context
8,192 tokens (as configured in the released config.json)
The base model was trained on 2 trillion tokens drawn from a mix of FineWeb-Edu, DCLM, The Stack
and further filtered datasets curated by the SmolLM team. The instruct version then went through
supervised fine-tuning on a combination of public and in-house datasets (smol-smoltalk), followed by
Direct Preference Optimization against UltraFeedback.
Good for: drafting, rewriting, summarisation, simple extraction, and as a test bed for tooling that
will later run a bigger model.
Do not rely on it for: factual accuracy, arithmetic beyond the trivial, long-context reasoning, or any
language other than English. A 135M model hallucinates confidently and has a shallow world model. It
has no safety fine-tuning beyond what the SFT mixture provided, so treat its output as untrusted text.