The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
A 256M-parameter vision-language model that reads video, not just stills, and needs about 1.4 GB of
GPU memory to do it. This is the smallest credible video-understanding model in the catalogue.
A multimodal model built on the Idefics3 architecture, taking interleaved text, images and video and
producing text. It answers questions about media, compares clips, and transcribes text it sees.
Parameters
256,484,928
File
model.safetensors, 1,025,998,224 bytes
Architecture
Idefics3 (SigLIP vision encoder + SmolLM2 language decoder)
Inputs
text, one or more images, video
Outputs
text only — it does not generate images or video
Languages
English
Licence
Apache-2.0
Video frames are encoded as a sequence of images, so the cost scales with the number of frames you feed
it, not with duration. That is the knob that decides whether this fits: a handful of sampled frames runs
on a laptop, a full-length clip at high frame rate does not.
Fine-tuned from HuggingFaceTB/SmolVLM-256M-Instruct on a mixture that includes The Cauldron, Docmatix,
LLaVA-OneVision-Data, LLaVA-Video-178K, ShareGPT4Video, Video-STaR, Vript and FineVideo. The video
instruction data is what separates this from its image-only predecessor.
The team reports modest but real numbers for a model this size (Video-MME 33.7, MLVU 40.6, MVBench 32.7
in the 256M row) and 1.38 GB of GPU RAM for video inference. Read those scores as "it can do the task,
badly" rather than as a comparison to a 7B video model.
Good for: describing a short clip, answering simple questions about video content, picking highlights
from footage, reading on-screen text, and fine-tuning on a narrow video domain where 256M is enough.
Do not rely on it for: temporal reasoning that spans more than a few sampled frames, precise counting,
fast motion, audio (it does not hear anything), or transcription of speech. It is a vision-language
model, not a speech model. At 256M it will confidently misread a scene, and its benchmark scores above
should temper any expectation of accuracy. It is English-only.