The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
Fifty layers, 25.6M parameters, 1,000 ImageNet classes. ResNet-50 is a decade old and still the first
thing to try when you need image classification that runs anywhere.
A 50-layer residual convolutional network. The residual connections — the identity shortcuts that let a
block learn a modification rather than a whole mapping — are the entire reason networks this deep became
trainable, and this is the checkpoint from the paper.
Parameters
25,610,152
File
model.safetensors, 102,482,854 bytes (fp32)
Input
RGB, 224 x 224
Output
1,000 ImageNet-1k class logits
Licence
Apache-2.0
At 100 MB it runs on a CPU at a few frames per second, which is why it survives in embedded systems and
in every benchmarking script written since 2015.
Supervised classification on ImageNet-1k — 1.28 million training images across 1,000 classes. This
is the port of the original torchvision weights into the Hugging Face transformers layout; the
optimiser settings and augmentations come from the ResNet paper's recipe.
from transformers import AutoImageProcessor, AutoModelForImageClassification
from PIL import Image
processor = AutoImageProcessor.from_pretrained("microsoft/resnet-50")
model = AutoModelForImageClassification.from_pretrained("microsoft/resnet-50")
image = Image.open("dog.jpg").convert("RGB")
inputs = processor(image, return_tensors="pt")
logits = model(**inputs).logits
print(model.config.id2label[logits.argmax(-1).item()])
If you only want a feature extractor, output_hidden_states=True and take the pooled output — a
2048-dimensional vector that is a decent general-purpose image descriptor.
Good for: a fast classification baseline, feature extraction, transfer learning to a new set of classes,
and sanity-checking your data pipeline.
Do not rely on it for: images containing anything outside ImageNet's 1,000 categories — it will pick
the nearest of those regardless, with no "none of the above". It does not localise objects, does not
detect anything (use DETR for that), and it is markedly weaker than modern
vision transformers on fine-grained or out-of-distribution images. Newer, smaller and better backbones
exist; this one is here because it is reliable and everyone knows what it does.