tinymodels
HuggingFaceTB

SmolVLM2-256M-Video-Instruct

A 256M-parameter vision-language model that watches video, not just stills.

video-text-to-text safetensors 256M params apache-2.0
Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/HuggingFaceTB/SmolVLM2-256M-Video-Instruct/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone clip of 16 frames at 224x224 + one questiongreedy decoding, transformers, float32, synthetic clip; one warm-up run then one timed run 2,841.4 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 3,171 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
Video-MMEthe 256M row of the card's evaluation table 33.7% not stated huggingface.co
MLVUsame table 40.6% not stated huggingface.co
MVBenchsame table 32.7% not stated huggingface.co

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

SmolVLM2-256M-Video-Instruct

A 256M-parameter vision-language model that reads video, not just stills, and needs about 1.4 GB of GPU memory to do it. This is the smallest credible video-understanding model in the catalogue.

What it is

A multimodal model built on the Idefics3 architecture, taking interleaved text, images and video and producing text. It answers questions about media, compares clips, and transcribes text it sees.

Parameters256,484,928
Filemodel.safetensors, 1,025,998,224 bytes
ArchitectureIdefics3 (SigLIP vision encoder + SmolLM2 language decoder)
Inputstext, one or more images, video
Outputstext only — it does not generate images or video
LanguagesEnglish
LicenceApache-2.0

Video frames are encoded as a sequence of images, so the cost scales with the number of frames you feed it, not with duration. That is the knob that decides whether this fits: a handful of sampled frames runs on a laptop, a full-length clip at high frame rate does not.

How it was trained

Fine-tuned from HuggingFaceTB/SmolVLM-256M-Instruct on a mixture that includes The Cauldron, Docmatix, LLaVA-OneVision-Data, LLaVA-Video-178K, ShareGPT4Video, Video-STaR, Vript and FineVideo. The video instruction data is what separates this from its image-only predecessor.

The team reports modest but real numbers for a model this size (Video-MME 33.7, MLVU 40.6, MVBench 32.7 in the 256M row) and 1.38 GB of GPU RAM for video inference. Read those scores as "it can do the task, badly" rather than as a comparison to a 7B video model.

Reference: SmolVLM2 blog, https://huggingface.co/blog/smolvlm2.

How to run it

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

path = "HuggingFaceTB/SmolVLM2-256M-Video-Instruct"
processor = AutoProcessor.from_pretrained(path)
model = AutoModelForImageTextToText.from_pretrained(path, torch_dtype=torch.bfloat16)

messages = [{
    "role": "user",
    "content": [
        {"type": "video", "path": "clip.mp4"},
        {"type": "text", "text": "What happens in this clip, in one sentence?"},
    ],
}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt",
).to(model.device, dtype=torch.bfloat16)

out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])

Intended use and limits

Good for: describing a short clip, answering simple questions about video content, picking highlights from footage, reading on-screen text, and fine-tuning on a narrow video domain where 256M is enough.

Do not rely on it for: temporal reasoning that spans more than a few sampled frames, precise counting, fast motion, audio (it does not hear anything), or transcription of speech. It is a vision-language model, not a speech model. At 256M it will confidently misread a scene, and its benchmark scores above should temper any expectation of accuracy. It is English-only.

Upstream