tinymodels
microsoft

Florence-2-base-ft

232M parameters covering captioning, OCR, detection and segmentation behind one prompt format.

Not mirrored here. This entry is a pointer: the download link hands you to the publisher at huggingface.co/microsoft/Florence-2-base-ft/resolve/main/model.safetensors. The expected sha256 travels with the redirect in the x-checksum-sha256 header. Why it works this way.
what was measuredresultmeasured onsource
measured hereone 640x480 photo (coco-val2017-000000039769.jpg) + '<OD>'greedy decoding, transformers, float32; one warm-up run then one timed run 2,124.4 ms Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
measured herepeak memory to load and run that inputpeak resident set of the measuring process, runtime included 2,109 MB Apple M4 Pro, 48 GB unified memory, macOS 15.6 scripts/measure.py
COCO Caption Karpathy test CIDErthe fine-tuned-model table in the card 140 CIDEr not stated huggingface.co
COCO detection val2017 mAPsame table 41.4 mAP not stated huggingface.co
VQAv2 test-dev accuracysame table 79.7% not stated huggingface.co

The rows marked measured here were run by this hub on Apple M4 Pro, 48 GB unified memory, macOS 15.6: one fixed input per task type, one warm-up run, one timed run, with the toolchain that model's format implies. Those are the only numbers on this page that can be compared with each other, and the milliseconds one is the column the catalogue shows. The state is honest, not a benchmark submission: a laptop, one thread of attention, no optimised serving stack, and one run.

The rest are the publisher's published results, quoted rather than re-run. Where a source does not say what the run was performed on, the column says so. A figure quoted in a different precision or on a different device is not the speed or the score you will get, and different units mean two published rows here are not comparable with each other.

The full measured table — every model, one fixed input each, the whole run cited as one dataset — is the measured bench dataset page.

Florence-2 base (fine-tuned)

232M parameters covering captioning, OCR, object detection, grounding and segmentation behind a single prompt interface. It is the most capable vision model that still fits comfortably in this catalogue.

What it is

A sequence-to-sequence vision foundation model from Microsoft: a DaViT vision encoder feeding a transformer encoder-decoder, driven by short text prompts instead of separate task heads. The same weights do many tasks depending only on what you ask for.

Parameters231,567,705 (0.23B)
Filemodel.safetensors, 463,221,266 bytes (fp16)
Prompt formattask tokens such as <CAPTION>, <OD>, <OCR>, <DENSE_REGION_CAPTION>
LicenceMIT
Precisiontrained and released in float16

This row is Florence-2-base-ft — the fine-tuned variant. The plain Florence-2-base is the pre-trained model before downstream task training; the -ft suffix matters, and the fine-tuned one is what you want for inference.

How it was trained

Pre-trained on FLD-5B: 5.4 billion annotations spanning 126 million images, covering captions, boxes, masks and referring expressions. The -ft variant then went through multi-task fine-tuning on a collection of downstream datasets. The prompt-based design means one set of weights learns all of it jointly rather than being re-headed per task.

Reference: Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks, https://arxiv.org/abs/2311.06242.

How to run it

import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "microsoft/Florence-2-base-ft", trust_remote_code=True, torch_dtype=torch.float16
)
processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base-ft", trust_remote_code=True)

image = Image.open("receipt.jpg").convert("RGB")
prompt = "<OCR>"   # or "<CAPTION>", "<OD>", "<DENSE_REGION_CAPTION>"
inputs = processor(text=prompt, images=image, return_tensors="pt")

generated = model.generate(
    input_ids=inputs["input_ids"], pixel_values=inputs["pixel_values"],
    max_new_tokens=1024, do_sample=False, num_beams=3,
)
decoded = processor.batch_decode(generated, skip_special_tokens=False)[0]
print(processor.post_process_generation(decoded, task=prompt, image_size=(image.width, image.height)))

trust_remote_code=True is required: the released weights rely on modelling code that ships with the repository rather than in a stable transformers release.

Intended use and limits

Good for: captioning photos, extracting text from screenshots and documents, finding objects, grounding phrases to regions, and doing all of that with one model instead of four.

Do not rely on it for: production accuracy without your own evaluation. OCR is strong on clean, upright text and weak on handwriting and rotated scans. Detection is class-agnostic in some prompt modes, and boxes are coarser than a dedicated detector's. Because it executes repository code, trust_remote_code means you are running code you have not read — check the revision you pin before you put this anywhere that matters. Its licence is MIT, which is friendlier than most vision foundation models, but the underlying FLD-5B data provenance is worth your own look.

Upstream