Models
GitHub
Discord
Docs
Pricing
Sign in
Download
Models
Download
GitHub
Discord
Docs
Pricing
Sign in
ahmadwaqar
/
smolvlm2-500m-video
:fp16
218
Downloads
Updated
1 month ago
Compact 500M vision-language model for video/image understanding. Supports visual QA, captioning, OCR, video analysis. Only 1.8GB VRAM. Built on SigLIP + SmolLM2. Available in Q8 and FP16. Apache 2.0 license.
Compact 500M vision-language model for video/image understanding. Supports visual QA, captioning, OCR, video analysis. Only 1.8GB VRAM. Built on SigLIP + SmolLM2. Available in Q8 and FP16. Apache 2.0 license.
Cancel
vision
smolvlm2-500m-video:fp16
...
/
params
53ed932be8fa · 57B
{
"num_ctx": 8192,
"stop": [
"<end_of_utterance>"
]
}