Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
Vision models · Ollama
Vision models on Ollama.
  • deepseek-v4.1-flash

    DeepSeek-V4.1-Flash is an advanced tool designed to enhance search capabilities, providing users with faster and more accurate results.

    vision tools thinking cloud

    39.8K  Pulls 1  Tag Updated  1 week ago

  • qwen3.8-flash-next

    This experimental preview of the architecture that will underpin Qwen4.

    vision tools thinking

    109.9K  Pulls 6  Tags Updated  2 weeks ago

  • glm-5.3-flash

    Z.ai's first natively multimodal model, approaching Claude Opus 4.8 on coding and agentic benchmarks with just 18B active parameters.

    vision tools thinking cloud

    145.2K  Pulls 1  Tag Updated  3 weeks ago

  • ornith-1.5

    Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.

    vision 9b 35b 397b

    329.1K  Pulls 3  Tags Updated  1 month ago

  • qwen3.8

    Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.

    vision tools thinking 27b

    2.3M  Pulls 12  Tags Updated  1 month ago

  • muse-glimmer

    Meta's latest open model built for always-on local agents. 30B parameters, licensed under Apache 2.0 and runs on a single GPU — tuned for tool use, long tasks, and failure recovery.

    vision tools thinking 30b

    219.2K  Pulls 15  Tags Updated  3 weeks ago

  • kimi-k3

    Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date.

    vision tools thinking cloud

    85.7K  Pulls 1  Tag Updated  1 month ago

  • kimi-k2.7-code

    Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built upon Kimi K2.6, with substantial improvements on real-world long-horizon coding tasks and roughly 30% lower thinking-token usage.

    vision tools thinking cloud

    240.2K  Pulls 1  Tag Updated  3 months ago

  • minimax-m3

    MiniMax M3: Coding & Agentic Frontier. 1M context window. Native Multimodality.

    vision tools thinking cloud

    512.3K  Pulls 1  Tag Updated  3 months ago

  • minicpm-v4.5

    A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

    vision 8b

    40.7K  Pulls 13  Tags Updated  3 months ago

  • minicpm-v4.6

    A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

    vision 1b

    51.6K  Pulls 13  Tags Updated  3 months ago

  • mistral-medium-3.5

    Mistral Medium 3.5 is the first flagship model of Mistral AI that merged instruction-following, reasoning, and coding in a single set of 128B weights.

    vision tools thinking 128b

    404.8K  Pulls 5  Tags Updated  4 months ago

  • nemotron3

    NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.

    vision tools thinking 33b

    663.5K  Pulls 4  Tags Updated  4 months ago

  • qwen3.6

    Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.

    vision tools thinking 27b 35b

    6.7M  Pulls 35  Tags Updated  2 weeks ago

  • medgemma1.5

    MedGemma 1.5 4B is an updated version of the MedGemma 4B model.

    vision 4b

    189.7K  Pulls 5  Tags Updated  5 months ago

  • medgemma

    MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension.

    vision 4b 27b

    402.1K  Pulls 9  Tags Updated  5 months ago

  • kimi-k2.6

    Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

    vision tools thinking cloud

    476.2K  Pulls 1  Tag Updated  5 months ago

  • gemma4

    Gemma 4 models are designed to deliver frontier-level performance at each size. They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.

    vision tools thinking audio cloud e2b e4b 12b 26b 31b

    25.4M  Pulls 50  Tags Updated  2 weeks ago

  • qwen3.5

    Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.

    vision tools thinking cloud 0.8b 2b 4b 9b 27b 35b 122b

    20.6M  Pulls 64  Tags Updated  2 weeks ago

  • glm-ocr

    GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.

    vision tools

    7.4M  Pulls 3  Tags Updated  7 months ago

© 2026 Ollama
Blog Support