tinymodels
distilbert

distilbert-base-uncased-finetuned-sst-2-english

A 67M-parameter sentiment classifier: positive or negative, distilled from BERT.

text-classification safetensors 67M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/distilbert/distilbert-base-uncased-finetuned-sst-2-english/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone fixed 24-word sentenceone sentence classified (one forward pass), transformers, float32; one warm-up run then one timed run 13.1 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 525 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
SST-2 validation accuracythe verified evaluation recorded in the model card 91.06% not stated huggingface.co
SST-2 validation F1same record 91.37% not stated huggingface.co
SST-2 validation AUCsame record 97.17% not stated huggingface.co

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

DistilBERT base uncased, fine-tuned on SST-2

A 67M-parameter binary sentiment classifier: give it English text, get back POSITIVE or NEGATIVE. It is the canonical "hello world" of text classification and still a fair baseline.

What it is

DistilBERT is BERT-base distilled down to 6 layers and 66,955,010 parameters — about 40% smaller and roughly 60% faster than BERT-base, retaining the large majority of its language understanding. This checkpoint is that student fine-tuned on the Stanford Sentiment Treebank.

Parameters66,955,010
Filemodel.safetensors, 267,832,558 bytes (fp32)
Layers / hidden size6 / 768
LabelsPOSITIVE, NEGATIVE
Maximum input512 word pieces (SST-2 trained on short sentences)
LicenceApache-2.0

How it was trained

Two stages. First, distillation: the 6-layer student is trained to match the output distribution of bert-base-uncased, which is where its language ability comes from. Second, fine-tuning on SST-2, the Stanford Sentiment Treebank — roughly 67,000 short movie-review sentences labelled positive or negative, part of the GLUE benchmark.

Reference: DistilBERT, a distilled version of BERT, https://arxiv.org/abs/1910.01108.

How to run it

from transformers import pipeline

classifier = pipeline("sentiment-analysis",
                      model="distilbert/distilbert-base-uncased-finetuned-sst-2-english")
print(classifier("The tiny model fit in memory and the whole thing took four seconds."))
# [{'label': 'POSITIVE', 'score': 0.999...}]

It is also a fine starting point for fine-tuning on your own labels:

from transformers import AutoTokenizer, AutoModelForSequenceClassification

tok = AutoTokenizer.from_pretrained("distilbert/distilbert-base-uncased-finetuned-sst-2-english")
model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert/distilbert-base-uncased-finetuned-sst-2-english", num_labels=2)

Intended use and limits

Good for: quick sentiment baselines, smoke-testing a classification pipeline, teaching, and as an initialisation for labelled data you actually have.

Do not rely on it for: anything with a third sentiment — this model has no NEUTRAL output and will force a verdict on neutral text. It is trained on movie-review sentences, so it degrades on long documents, on modern slang, on technical text, and on sarcasm. It is uncased, so it cannot use capitalisation as evidence. Anything where a wrong label is costly deserves a model you fine-tuned on your own distribution.

Upstream