64 2 weeks ago

SmolVLM2-2.2B-Instruct is a compact multimodal model for image and video understanding. Built on SmolLM2-1.7B with SigLIP vision encoder. Supports visual QA, OCR, and video analysis. Available in Q8 and FP16 quantizations. Apache 2.0 license.

vision
ff1c370ef838 · 80B
{
"num_ctx": 8192,
"stop": [
"<|im_end|>",
"<end_of_utterance>"
]
}