The most powerful vision-language model in the Qwen model family to date.
5.9M Pulls 57 Tags Updated 10 months ago
DeepSeek-OCR is a vision-language model that can perform token-efficient OCR.
529.2K Pulls 3 Tags Updated 9 months ago
Flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL.
4.8M Pulls 17 Tags Updated 1 year ago
🌋 LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.
14.8M Pulls 98 Tags Updated 2 years ago
Llama 3.2 Vision is a collection of instruction-tuned image reasoning generative models in 11B and 90B sizes.
5.2M Pulls 9 Tags Updated 1 year ago
A series of multimodal LLMs (MLLMs) designed for vision-language understanding.
5.5M Pulls 17 Tags Updated 1 year ago
A compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more.
995K Pulls 5 Tags Updated 1 year ago
moondream2 is a small vision language model designed to run efficiently on edge devices.
1.8M Pulls 18 Tags Updated 2 years ago
Building upon Mistral Small 3, Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance.
786.5K Pulls 5 Tags Updated 1 year ago
MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension.
362.5K Pulls 9 Tags Updated 4 months ago
29 Pulls 1 Tag Updated 1 year ago
A family of open-source models trained on a wide variety of data, surpassing ChatGPT on various benchmarks. Updated to version 3.5-0106.
1.3M Pulls 50 Tags Updated 2 years ago
Text + Vision Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-NVFP4
690 Pulls 1 Tag Updated 4 days ago
v4's merge rebased on Qwen/Qwen3.8-27B (+ 3 Qwen3.6 fine-tunes) — +6.5pp LiveCodeBench, MTP + vision
461 Pulls 39 Tags Updated 9 minutes ago
gemma4:31b-mtp-vision-bf16
7 Pulls 1 Tag Updated 3 days ago
Qwen3.8-27B tensor-level abliterated, vision tower and MTP head untouched. 0% over-refusal on XSTest, 0-6% refusal across the A/B suite, no measurable capability loss. Full mmproj vision, tool calling and thinking, 262K context.
145K Pulls 20 Tags Updated 1 week ago
5,838 Pulls 4 Tags Updated 1 week ago
😈 Uncensored Qwen3.8-27B Vision (27.3B • Q4_K_M) for Ollama. 👁️ Multimodal Vision AI with CLIP-ViT, Thinking, MTP, Tool Calling, Rust 1.98.0 Image Analysis, GGUF & local AI development. Medic🚀 #16HEX Matrix / Nightshift Heretic.
2,460 Pulls 1 Tag Updated 2 weeks ago
Text + Vision Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF-Q8-NVFP4
830 Pulls 1 Tag Updated 2 weeks ago
Ollama repack of Qwen3.6-35B-A3B Genesis Hermes V7 with APEX/APEX-Compact, Vision, and 224K context for coding agents.
859 Pulls 3 Tags Updated 3 weeks ago