The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
VideoMAE extends masked autoencoders to video. The architecture is close to a plain vision transformer
over spatio-temporal patches — 16x16 pixels per patch, 16 frames per clip — with a light decoder used
during pre-training to reconstruct masked patches. For classification, a linear head sits on the [CLS]
token.
Parameters
86,534,800
File
model.safetensors, 346,161,616 bytes (fp32)
Input
16 frames at 224 x 224
Output
400 Kinetics-400 action classes
Licence
CC BY-NC 4.0 — non-commercial use only
The non-commercial licence is the first thing to check here. Kinetics-400 is derived from YouTube video,
and the model inherits a restriction that most of the models in this catalogue do not have. For anything
that ships commercially, pick a different backbone.
Self-supervised pre-training for 1,600 epochs on Kinetics-400 video without labels — the model
learns by reconstructing heavily masked patches, which is the paper's central claim: masked video
modelling is data-efficient. A supervised fine-tuning pass on the labelled Kinetics-400 split then adds
the classification head.
The model card upstream was written by Hugging Face, not by the paper's authors; the training-data and
preprocessing sections there are still marked "to do".
Reference: VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video
Pre-Training, https://arxiv.org/abs/2203.12602.
import numpy as np
import torch
from transformers import VideoMAEImageProcessor, VideoMAEForVideoClassification
processor = VideoMAEImageProcessor.from_pretrained("MCG-NJU/videomae-base-finetuned-kinetics")
model = VideoMAEForVideoClassification.from_pretrained("MCG-NJU/videomae-base-finetuned-kinetics")
# 16 frames, 3 channels, 224x224, in [0, 255] as uint8
video = list(np.random.randint(0, 255, (16, 3, 224, 224)).astype("uint8"))
inputs = processor(video, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
print(model.config.id2label[logits.argmax(-1).item()])
Sampling the 16 frames evenly across your clip matters more than anything else about how you call this.
Too dense and a slow action looks static; too sparse and you skip the action entirely.
Good for: labelling short trimmed clips into coarse human actions, experimenting with video
transformers, and pre-training research on masked video modelling.
Do not rely on it for: commercial products (the licence forbids it), untrimmed video where the action is
a small part of the clip, fine-grained distinctions between similar actions, or anything with audio. It
sees 16 frames and no motion vectors, so fast actions blur into nothing. Kinetics-400's 400 classes are
a coarse vocabulary — most of what happens in real footage is not in it.