The rows marked measured here were run by this hub on
Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task
type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only
numbers on this page that can be compared with each other, and the milliseconds one is the column the
catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no
optimised serving stack, and one run.
The rest are the publisher's published results, quoted rather than re-run. Where a source does not say
what the run was performed on, the column says so. A figure quoted in a different precision or on a
different device is not the speed or the score you will get, and different units mean two published rows
here are not comparable with each other.
The full measured table — every model, one fixed input each, the whole run cited as one dataset — is
the measured bench dataset page.
232M parameters covering captioning, OCR, object detection, grounding and segmentation behind a single
prompt interface. It is the most capable vision model that still fits comfortably in this catalogue.
A sequence-to-sequence vision foundation model from Microsoft: a DaViT vision encoder feeding a
transformer encoder-decoder, driven by short text prompts instead of separate task heads. The same
weights do many tasks depending only on what you ask for.
Parameters
231,567,705 (0.23B)
File
model.safetensors, 463,221,266 bytes (fp16)
Prompt format
task tokens such as <CAPTION>, <OD>, <OCR>, <DENSE_REGION_CAPTION>
Licence
MIT
Precision
trained and released in float16
This row is Florence-2-base-ft — the fine-tuned variant. The plain Florence-2-base is the
pre-trained model before downstream task training; the -ft suffix matters, and the fine-tuned one is
what you want for inference.
Pre-trained on FLD-5B: 5.4 billion annotations spanning 126 million images, covering captions,
boxes, masks and referring expressions. The -ft variant then went through multi-task fine-tuning on a
collection of downstream datasets. The prompt-based design means one set of weights learns all of it
jointly rather than being re-headed per task.
trust_remote_code=True is required: the released weights rely on modelling code that ships with the
repository rather than in a stable transformers release.
Good for: captioning photos, extracting text from screenshots and documents, finding objects, grounding
phrases to regions, and doing all of that with one model instead of four.
Do not rely on it for: production accuracy without your own evaluation. OCR is strong on clean, upright
text and weak on handwriting and rotated scans. Detection is class-agnostic in some prompt modes, and
boxes are coarser than a dedicated detector's. Because it executes repository code, trust_remote_code
means you are running code you have not read — check the revision you pin before you put this anywhere
that matters. Its licence is MIT, which is friendlier than most vision foundation models, but the
underlying FLD-5B data provenance is worth your own look.