tinymodels
tasks · tinymodels.co

Tiny video-text-to-text models

Every video-text-to-text model in the catalogue, with the size and licence stated and one measured millisecond figure per model — one fixed input, one machine, one method.

Browse by format instead: gguf, mlx, safetensors.

modelparamssizelicence one fixed input (ms)
SmolVLM2-256M-Video-Instruct 256M 1.03 GB apache-2.0 2,841.4 Get

The millisecond figure is one fixed input per task type, measured on this hub's machine with one warm-up run then one timed run. How these numbers were taken; the whole run, every model, is the measured bench dataset.