tasks · tinymodels.co
Tiny video-text-to-text models
Every video-text-to-text model in the catalogue, with the size and licence stated and one measured millisecond figure per model — one fixed input, one machine, one method.
Browse by format instead: gguf, mlx, safetensors.
| model | params | size | licence | one fixed input (ms) | |
|---|---|---|---|---|---|
| SmolVLM2-256M-Video-Instruct | 256M | 1.03 GB | apache-2.0 | 2,841.4 | Get |
The millisecond figure is one fixed input per task type, measured on this hub's machine with one warm-up run then one timed run. How these numbers were taken; the whole run, every model, is the measured bench dataset.