The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
A 22.7M-parameter sentence encoder that turns a sentence into a 384-dimensional vector. It is one of
the most-downloaded models on the internet for a simple reason: it is small, fast, and good enough for
semantic search on a laptop.
A MiniLM-style transformer: 6 layers, 384 hidden dimensions, 22,713,728 parameters, 90 MB of
safetensors. It maps a sentence or short paragraph to a fixed-length embedding such that
semantically similar text lands nearby in vector space.
Distilled from a larger teacher into a 6-layer MiniLM student, then fine-tuned on a 1 billion
sentence-pair corpus built from sources including S2ORC, StackExchange, MS MARCO, GooAQ, Natural
Questions, TriviaQA, SNLI, MultiNLI, WikiHow and a set of NLI-based paraphrase collections. The
published recipe is on the sentence-transformers model card; nreimers/MiniLM-L6-H384-uncased is the
intermediate checkpoint this was built from.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
sentences = [
"A tiny model that runs on a laptop.",
"Small models you can run locally.",
"The price of tea in Lisbon.",
]
embeddings = model.encode(sentences, normalize_embeddings=True)
print(embeddings.shape) # (3, 384)
print((embeddings[0] @ embeddings[1])) # high
print((embeddings[0] @ embeddings[2])) # much lower
The repository also ships ONNX and OpenVINO builds, which is why this model shows up inside local
search features and browser extensions that have no Python in them.
Good for: semantic search over a few thousand documents, clustering, deduplication, topic tagging,
retrieval as the first stage of a RAG pipeline, and as a cheap similarity baseline.
Do not rely on it for: anything past 256 word pieces — this is the single most common way to get bad
results from it, because the tail of a long document is silently dropped. It is English-only in
practice, is not a reranker, and does not handle long documents, code or numeric reasoning well. For
higher quality at 5x the size, BAAI/bge-base-en-v1.5 and the all-mpnet-base-v2 model are the usual
upgrades.