The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
DEtection TRansformer pairs a ResNet-50 convolutional backbone with a transformer encoder-decoder. The
decoder emits a fixed set of predictions and uses bipartite matching against the ground truth during
training, which is what removes the anchors and the NMS post-processing that every other detector
needs.
Supervised on COCO 2017 — 118,000 annotated images — with the standard DETR recipe. The Hungarian
matching loss assigns each prediction to a ground-truth box, and the "no object" class absorbs the
predictions that match nothing.
Good for: clean, well-framed objects at moderate scale; a reference implementation of detection without
post-processing; and describing what is in a photo well enough to index it.
Do not rely on it for: small objects, dense crowds, or real-time video. Training on 118k images means it
generalises worse than modern detectors trained on far more, and it is notably slow — a full forward
pass over a large image takes hundreds of milliseconds on CPU. Detecting text or reading plates needs a
specialised model, not this. See Florence-2 in this catalogue for a model
that handles several of these tasks at once.