The most powerful vision-language model in the Qwen model family to date.
4.7M Pulls 57 Tags Updated 8 months ago
Kimi K2.5 is an open-source, native multimodal agentic model that seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.
368.2K Pulls 1 Tag Updated 5 months ago
DeepSeek-OCR is a vision-language model that can perform token-efficient OCR.
497.3K Pulls 3 Tags Updated 8 months ago
🌋 LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.
14.4M Pulls 98 Tags Updated 2 years ago
Flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL.
3.5M Pulls 17 Tags Updated 1 year ago
Llama 3.2 Vision is a collection of instruction-tuned image reasoning generative models in 11B and 90B sizes.
4.9M Pulls 9 Tags Updated 1 year ago
A series of multimodal LLMs (MLLMs) designed for vision-language understanding.
5.3M Pulls 17 Tags Updated 1 year ago
A compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more.
950.8K Pulls 5 Tags Updated 1 year ago
Building upon Mistral Small 3, Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance.
766.3K Pulls 5 Tags Updated 1 year ago
moondream2 is a small vision language model designed to run efficiently on edge devices.
1.5M Pulls 18 Tags Updated 2 years ago
MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension.
201.9K Pulls 9 Tags Updated 3 months ago
29 Pulls 1 Tag Updated 10 months ago
A family of open-source models trained on a wide variety of data, surpassing ChatGPT on various benchmarks. Updated to version 3.5-0106.
1.2M Pulls 50 Tags Updated 2 years ago
NuExtract3 is a 4B vision-language reasoning model by NuMind for local document understanding. Turn text and document images into structured JSON or clean Markdown with Ollama, with support for multilingual documents, OCR, and template generation.
118 Pulls 4 Tags Updated 3 days ago
Gemma 4 31B multimodal instruct model with vision support, quantized to Q3_K_S and optimized for 16 GB VRAM. Supports a practical context window of approximately 36K–40K tokens, depending on the Ollama version, GPU, backend and runtime configuration.
84 Pulls 1 Tag Updated 6 days ago
Agentic coding model for 24 GB GPUs: TeichAI's Gemma-4-31B Fable-5 agent distill (vision + thinking + tools) with a disciplined coding-agent system prompt baked in. Inspect → reproduce → smallest fix → re-verify.
66 Pulls 1 Tag Updated 6 days ago
Qwen3.5 2B in Q8_0 quantization. Strong balance of capability and efficiency with 262K context, vision, tool use, and thinking. Ideal for local deployment on consumer hardware.
37 Pulls 1 Tag Updated 3 days ago
Qwen3.5 0.8B in Q8_0 quantization. Small, fast model with 262K context, vision, tool use, and thinking capabilities. Optimized for local/edge deployment on constrained hardware.
21 Pulls 1 Tag Updated 3 days ago
Ollama Agent + Skills + Vision 26 t/s en 8 VRAM 2026
5 Pulls 1 Tag Updated 2 days ago
MedGemma 1.5 4.3B Q4_K_M with native Ollama thinking separation, vision support, deterministic decoding, and unchanged official weights and tokenizer.
4 Pulls 1 Tag Updated 19 hours ago