The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
A 67M-parameter binary sentiment classifier: give it English text, get back POSITIVE or NEGATIVE.
It is the canonical "hello world" of text classification and still a fair baseline.
DistilBERT is BERT-base distilled down to 6 layers and 66,955,010 parameters — about 40% smaller and
roughly 60% faster than BERT-base, retaining the large majority of its language understanding. This
checkpoint is that student fine-tuned on the Stanford Sentiment Treebank.
Parameters
66,955,010
File
model.safetensors, 267,832,558 bytes (fp32)
Layers / hidden size
6 / 768
Labels
POSITIVE, NEGATIVE
Maximum input
512 word pieces (SST-2 trained on short sentences)
Two stages. First, distillation: the 6-layer student is trained to match the output distribution of
bert-base-uncased, which is where its language ability comes from. Second, fine-tuning on SST-2, the
Stanford Sentiment Treebank — roughly 67,000 short movie-review sentences labelled positive or negative,
part of the GLUE benchmark.
from transformers import pipeline
classifier = pipeline("sentiment-analysis",
model="distilbert/distilbert-base-uncased-finetuned-sst-2-english")
print(classifier("The tiny model fit in memory and the whole thing took four seconds."))
# [{'label': 'POSITIVE', 'score': 0.999...}]
It is also a fine starting point for fine-tuning on your own labels:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tok = AutoTokenizer.from_pretrained("distilbert/distilbert-base-uncased-finetuned-sst-2-english")
model = AutoModelForSequenceClassification.from_pretrained(
"distilbert/distilbert-base-uncased-finetuned-sst-2-english", num_labels=2)
Good for: quick sentiment baselines, smoke-testing a classification pipeline, teaching, and as an
initialisation for labelled data you actually have.
Do not rely on it for: anything with a third sentiment — this model has no NEUTRAL output and will
force a verdict on neutral text. It is trained on movie-review sentences, so it degrades on long
documents, on modern slang, on technical text, and on sarcasm. It is uncased, so it cannot use
capitalisation as evidence. Anything where a wrong label is costly deserves a model you fine-tuned on
your own distribution.