The most powerful vision-language model in the Qwen model family to date.
5.1M Pulls 57 Tags Updated 9 months ago
DeepSeek-OCR is a vision-language model that can perform token-efficient OCR.
511.4K Pulls 3 Tags Updated 8 months ago
🌋 LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.
14.6M Pulls 98 Tags Updated 2 years ago
Flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL.
4.1M Pulls 17 Tags Updated 1 year ago
Llama 3.2 Vision is a collection of instruction-tuned image reasoning generative models in 11B and 90B sizes.
5M Pulls 9 Tags Updated 1 year ago
A series of multimodal LLMs (MLLMs) designed for vision-language understanding.
5.4M Pulls 17 Tags Updated 1 year ago
A compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more.
971.3K Pulls 5 Tags Updated 1 year ago
Building upon Mistral Small 3, Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance.
777K Pulls 5 Tags Updated 1 year ago
moondream2 is a small vision language model designed to run efficiently on edge devices.
1.5M Pulls 18 Tags Updated 2 years ago
MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension.
266.1K Pulls 9 Tags Updated 3 months ago
29 Pulls 1 Tag Updated 11 months ago
A family of open-source models trained on a wide variety of data, surpassing ChatGPT on various benchmarks. Updated to version 3.5-0106.
1.3M Pulls 50 Tags Updated 2 years ago
a 27.9B vision-language model tuned for agentic workflows.
1 Pull 1 Tag Updated 15 hours ago
1,813 Pulls 1 Tag Updated 1 week ago
NuExtract3 is a 4B vision-language reasoning model by NuMind for local document understanding. Turn text and document images into structured JSON or clean Markdown with Ollama, with support for multilingual documents, OCR, and template generation.
787 Pulls 4 Tags Updated 3 weeks ago
To run DeepSeek-V4-Flash in full precision lossless, run Q4 (UD-Q4_K_XL), It is 155GB. vision tools thinking
533 Pulls 1 Tag Updated 2 weeks ago
Next-Gen Sovereign AI Ecosystem - 7 specialized models for Coding, Vision, Reasoning, Edge & RAG. 131K context, 11 Stop Tokens. Native with OpenClaw, VSCode, Cursor, LangChain & 12+ platforms. Forged by NJIRLAH.
2.3M Pulls 7 Tags Updated 2 months ago
Multimodal 4.5B Gemma 4 with Unsloth Dynamic QAT & MTP (5.3GB). Features near-lossless vision, audio, text reasoning, an active 4-token speculative window, and an expanded 64K context size. Best for edge execution on systems with 8GB to 16GB RAM.
592 Pulls 1 Tag Updated 3 weeks ago
Ultra-fast multimodal 2.3B Gemma 4 for on-device edge AI (3.7GB). Adds native vision/audio parsing to Unsloth Dynamic QAT with a 2-token MTP pipeline and a mobile-safe 32K context window. Perfect for smartphones and laptops sharing 8GB of total RAM.
475 Pulls 1 Tag Updated 3 weeks ago
Qwen 3.6 35B multimodal model with vision support, quantized to IQ3_S and optimized to fit within 16 GB VRAM with Q4_0 KV cache. Supports very large context windows, with the practical maximum depending on the Ollama version, GPU and backend.
423 Pulls 1 Tag Updated 3 weeks ago