55 2 weeks ago

Compact 500M vision-language model for video/image understanding. Supports visual QA, captioning, OCR, video analysis. Only 1.8GB VRAM. Built on SigLIP + SmolLM2. Available in Q8 and FP16. Apache 2.0 license.

vision
79c05ad2d80e · 64B
Apache 2.0 License - https://www.apache.org/licenses/LICENSE-2.0