tinymodels
microsoft

resnet-50

ResNet-50: 25M parameters, 1000 ImageNet classes, the baseline that still refuses to die.

image-classification safetensors 25.6M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/microsoft/resnet-50/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 640x480 photo (coco-val2017-000000039769.jpg)one forward pass, transformers, float32; one warm-up run then one timed run 21.7 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 505 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
ImageNet-1k top-1, 224pxtimm resnet50.a1_in1k, the ImageNet-1k weights this checkpoint was converted from 80.38% not stated github.com
ImageNet-1k top-5, 224pxsame results file 94.6% not stated github.com

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

ResNet-50

Fifty layers, 25.6M parameters, 1,000 ImageNet classes. ResNet-50 is a decade old and still the first thing to try when you need image classification that runs anywhere.

What it is

A 50-layer residual convolutional network. The residual connections — the identity shortcuts that let a block learn a modification rather than a whole mapping — are the entire reason networks this deep became trainable, and this is the checkpoint from the paper.

Parameters25,610,152
Filemodel.safetensors, 102,482,854 bytes (fp32)
InputRGB, 224 x 224
Output1,000 ImageNet-1k class logits
LicenceApache-2.0

At 100 MB it runs on a CPU at a few frames per second, which is why it survives in embedded systems and in every benchmarking script written since 2015.

How it was trained

Supervised classification on ImageNet-1k — 1.28 million training images across 1,000 classes. This is the port of the original torchvision weights into the Hugging Face transformers layout; the optimiser settings and augmentations come from the ResNet paper's recipe.

Reference: Deep Residual Learning for Image Recognition, https://arxiv.org/abs/1512.03385.

How to run it

from transformers import AutoImageProcessor, AutoModelForImageClassification
from PIL import Image

processor = AutoImageProcessor.from_pretrained("microsoft/resnet-50")
model = AutoModelForImageClassification.from_pretrained("microsoft/resnet-50")

image = Image.open("dog.jpg").convert("RGB")
inputs = processor(image, return_tensors="pt")
logits = model(**inputs).logits
print(model.config.id2label[logits.argmax(-1).item()])

If you only want a feature extractor, output_hidden_states=True and take the pooled output — a 2048-dimensional vector that is a decent general-purpose image descriptor.

Intended use and limits

Good for: a fast classification baseline, feature extraction, transfer learning to a new set of classes, and sanity-checking your data pipeline.

Do not rely on it for: images containing anything outside ImageNet's 1,000 categories — it will pick the nearest of those regardless, with no "none of the above". It does not localise objects, does not detect anything (use DETR for that), and it is markedly weaker than modern vision transformers on fine-grained or out-of-distribution images. Newer, smaller and better backbones exist; this one is here because it is reliable and everyone knows what it does.

Upstream