tinymodels
facebook

detr-resnet-50

DETR with a ResNet-50 backbone: 42M parameters of end-to-end object detection.

object-detection safetensors 41.6M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/facebook/detr-resnet-50/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 640x480 photo (coco-val2017-000000039769.jpg)one forward pass, transformers, float32; one warm-up run then one timed run 251.1 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 1,373 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
COCO 2017 validation APthe figure the model card states for this checkpoint 42 AP not stated huggingface.co

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

DETR ResNet-50

End-to-end object detection at 42M parameters: no anchor boxes, no non-maximum suppression, just a transformer that predicts a set of objects.

What it is

DEtection TRansformer pairs a ResNet-50 convolutional backbone with a transformer encoder-decoder. The decoder emits a fixed set of predictions and uses bipartite matching against the ground truth during training, which is what removes the anchors and the NMS post-processing that every other detector needs.

Parameters41,631,008
Filemodel.safetensors, 166,587,896 bytes (fp32)
BackboneResNet-50
OutputCOCO class boxes, 91 categories + no-object
LicenceApache-2.0

How it was trained

Supervised on COCO 2017 — 118,000 annotated images — with the standard DETR recipe. The Hungarian matching loss assigns each prediction to a ground-truth box, and the "no object" class absorbs the predictions that match nothing.

Reference: End-to-End Object Detection with Transformers, https://arxiv.org/abs/2005.12872.

How to run it

from transformers import AutoImageProcessor, DetrForObjectDetection
from PIL import Image
import torch

processor = AutoImageProcessor.from_pretrained("facebook/detr-resnet-50")
model = DetrForObjectDetection.from_pretrained("facebook/detr-resnet-50")

image = Image.open("street.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

target_sizes = torch.tensor([image.size[::-1]])
results = processor.post_process_object_detection(outputs, target_sizes=target_sizes, threshold=0.9)[0]
for score, label, box in zip(results["scores"], results["labels"], results["boxes"]):
    print(model.config.id2label[label.item()], round(score.item(), 3), [round(v, 1) for v in box.tolist()])

threshold=0.9 is doing real work there — DETR's small-object behaviour is its weak point, and a low threshold floods you with spurious boxes.

Intended use and limits

Good for: clean, well-framed objects at moderate scale; a reference implementation of detection without post-processing; and describing what is in a photo well enough to index it.

Do not rely on it for: small objects, dense crowds, or real-time video. Training on 118k images means it generalises worse than modern detectors trained on far more, and it is notably slow — a full forward pass over a large image takes hundreds of milliseconds on CPU. Detecting text or reading plates needs a specialised model, not this. See Florence-2 in this catalogue for a model that handles several of these tasks at once.

Upstream