tinymodels
sentence-transformers

all-MiniLM-L6-v2

The 22M-parameter sentence embedder that quietly powers a large share of local search.

sentence-similarity safetensors 22.7M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone fixed 24-word sentenceone sentence embedded (one forward pass), transformers, float32; one warm-up run then one timed run 5.3 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 395 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
Encoding speedthe sentence-transformers model table's own definition of its speed column 14,200 sentences/s NVIDIA V100 GPU sbert.net

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

all-MiniLM-L6-v2

A 22.7M-parameter sentence encoder that turns a sentence into a 384-dimensional vector. It is one of the most-downloaded models on the internet for a simple reason: it is small, fast, and good enough for semantic search on a laptop.

What it is

A MiniLM-style transformer: 6 layers, 384 hidden dimensions, 22,713,728 parameters, 90 MB of safetensors. It maps a sentence or short paragraph to a fixed-length embedding such that semantically similar text lands nearby in vector space.

Parameters22,713,728
Filemodel.safetensors, 90,868,376 bytes
Layers / hidden size6 / 384
Embedding dimension384
Maximum input256 word pieces (longer input is truncated)
LicenceApache-2.0

How it was trained

Distilled from a larger teacher into a 6-layer MiniLM student, then fine-tuned on a 1 billion sentence-pair corpus built from sources including S2ORC, StackExchange, MS MARCO, GooAQ, Natural Questions, TriviaQA, SNLI, MultiNLI, WikiHow and a set of NLI-based paraphrase collections. The published recipe is on the sentence-transformers model card; nreimers/MiniLM-L6-H384-uncased is the intermediate checkpoint this was built from.

How to run it

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
sentences = [
    "A tiny model that runs on a laptop.",
    "Small models you can run locally.",
    "The price of tea in Lisbon.",
]
embeddings = model.encode(sentences, normalize_embeddings=True)

print(embeddings.shape)          # (3, 384)
print((embeddings[0] @ embeddings[1]))  # high
print((embeddings[0] @ embeddings[2]))  # much lower

The repository also ships ONNX and OpenVINO builds, which is why this model shows up inside local search features and browser extensions that have no Python in them.

Intended use and limits

Good for: semantic search over a few thousand documents, clustering, deduplication, topic tagging, retrieval as the first stage of a RAG pipeline, and as a cheap similarity baseline.

Do not rely on it for: anything past 256 word pieces — this is the single most common way to get bad results from it, because the tail of a long document is silently dropped. It is English-only in practice, is not a reranker, and does not handle long documents, code or numeric reasoning well. For higher quality at 5x the size, BAAI/bge-base-en-v1.5 and the all-mpnet-base-v2 model are the usual upgrades.

Upstream