tinymodels
HuggingFaceTB

SmolLM2-135M-Instruct

A 135M-parameter instruction-tuned chat model small enough to run in a browser tab.

text-generation safetensors 135M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 64-token reply to a fixed promptgreedy decoding, mlx-lm on the GPU, publisher bf16 weights; one warm-up run then one timed run 110.2 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 658 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
HellaSwag (0-shot)instruction-model table in the model card, scored with lighteval 40.9% not stated huggingface.co
MMLU (cloze, 0-shot)same table 29.3% not stated huggingface.co
ARC (average, 0-shot)same table 37.3% not stated huggingface.co
BBH (3-shot)same table 28.2% not stated huggingface.co

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

SmolLM2-135M-Instruct

A 135M-parameter instruction-tuned language model from Hugging Face's SmolLM2 family, built for on-device chat where a laptop or a browser tab has to do the work.

What it is

SmolLM2 is a compact model family at 135M, 360M and 1.7B parameters. This is the smallest instruct variant: a Llama-style decoder-only transformer with 134,515,008 parameters, stored here as a 269 MB safetensors file in bf16.

Parameters134,515,008
Filemodel.safetensors, 269,060,552 bytes
ArchitectureLlama-style decoder-only transformer
Context8,192 tokens (as configured in the released config.json)
LanguagesEnglish
LicenceApache-2.0

How it was trained

The base model was trained on 2 trillion tokens drawn from a mix of FineWeb-Edu, DCLM, The Stack and further filtered datasets curated by the SmolLM team. The instruct version then went through supervised fine-tuning on a combination of public and in-house datasets (smol-smoltalk), followed by Direct Preference Optimization against UltraFeedback.

Reference: SmolLM2: When Smol Goes Big — Data-Centric Training of a Small Language Model, https://arxiv.org/abs/2502.02737.

How to run it

from transformers import AutoModelForCausalLM, AutoTokenizer

checkpoint = "HuggingFaceTB/SmolLM2-135M-Instruct"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint)  # device="cpu" is genuinely fine here

messages = [{"role": "user", "content": "What is gravity?"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False)
inputs = tokenizer.encode(prompt, return_tensors="pt")
outputs = model.generate(inputs, max_new_tokens=64, temperature=0.2, top_p=0.9, do_sample=True)
print(tokenizer.decode(outputs[0]))

There is a chat CLI if you would rather not write code:

pip install trl
trl chat --model_name_or_path HuggingFaceTB/SmolLM2-135M-Instruct --device cpu

The upstream repository also ships ONNX and Transformers.js weights, which is how this model ends up running inside a web page.

Intended use and limits

Good for: drafting, rewriting, summarisation, simple extraction, and as a test bed for tooling that will later run a bigger model.

Do not rely on it for: factual accuracy, arithmetic beyond the trivial, long-context reasoning, or any language other than English. A 135M model hallucinates confidently and has a shallow world model. It has no safety fine-tuning beyond what the SFT mixture provided, so treat its output as untrusted text.

Upstream