tinymodels
MCG-NJU

videomae-base-finetuned-kinetics

VideoMAE-base at 87M parameters, labelling 400 human actions from a clip.

video-classification safetensors 86.5M params cc-by-nc-4.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/MCG-NJU/videomae-base-finetuned-kinetics/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone clip of 16 frames at 224x224one forward pass, transformers, float32, synthetic clip; one warm-up run then one timed run 214.2 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 976 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
Kinetics-400 test top-1the evaluation results the model card states 80.9% not stated huggingface.co
Kinetics-400 test top-5same record 94.7% not stated huggingface.co

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

VideoMAE base, fine-tuned on Kinetics-400

An 87M-parameter video classifier: hand it a clip, get back one of 400 human actions. Note the licence — this one is non-commercial.

What it is

VideoMAE extends masked autoencoders to video. The architecture is close to a plain vision transformer over spatio-temporal patches — 16x16 pixels per patch, 16 frames per clip — with a light decoder used during pre-training to reconstruct masked patches. For classification, a linear head sits on the [CLS] token.

Parameters86,534,800
Filemodel.safetensors, 346,161,616 bytes (fp32)
Input16 frames at 224 x 224
Output400 Kinetics-400 action classes
LicenceCC BY-NC 4.0 — non-commercial use only

The non-commercial licence is the first thing to check here. Kinetics-400 is derived from YouTube video, and the model inherits a restriction that most of the models in this catalogue do not have. For anything that ships commercially, pick a different backbone.

How it was trained

Self-supervised pre-training for 1,600 epochs on Kinetics-400 video without labels — the model learns by reconstructing heavily masked patches, which is the paper's central claim: masked video modelling is data-efficient. A supervised fine-tuning pass on the labelled Kinetics-400 split then adds the classification head.

The model card upstream was written by Hugging Face, not by the paper's authors; the training-data and preprocessing sections there are still marked "to do".

Reference: VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, https://arxiv.org/abs/2203.12602.

How to run it

import numpy as np
import torch
from transformers import VideoMAEImageProcessor, VideoMAEForVideoClassification

processor = VideoMAEImageProcessor.from_pretrained("MCG-NJU/videomae-base-finetuned-kinetics")
model = VideoMAEForVideoClassification.from_pretrained("MCG-NJU/videomae-base-finetuned-kinetics")

# 16 frames, 3 channels, 224x224, in [0, 255] as uint8
video = list(np.random.randint(0, 255, (16, 3, 224, 224)).astype("uint8"))
inputs = processor(video, return_tensors="pt")
with torch.no_grad():
    logits = model(**inputs).logits
print(model.config.id2label[logits.argmax(-1).item()])

Sampling the 16 frames evenly across your clip matters more than anything else about how you call this. Too dense and a slow action looks static; too sparse and you skip the action entirely.

Intended use and limits

Good for: labelling short trimmed clips into coarse human actions, experimenting with video transformers, and pre-training research on masked video modelling.

Do not rely on it for: commercial products (the licence forbids it), untrimmed video where the action is a small part of the clip, fine-grained distinctions between similar actions, or anything with audio. It sees 16 frames and no motion vectors, so fast actions blur into nothing. Kinetics-400's 400 classes are a coarse vocabulary — most of what happens in real footage is not in it.

Upstream