Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
LLaVA · Ollama
Search for models on Ollama.
  • llava

    🌋 LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.

    vision 7b 13b 34b

    15M  Pulls 98  Tags Updated  2 years ago

  • llava-llama3

    A LLaVA model fine-tuned from Llama 3 Instruct with better scores in several benchmarks.

    vision 8b

    2.3M  Pulls 4  Tags Updated  2 years ago

  • bakllava

    BakLLaVA is a multimodal model consisting of the Mistral 7B base model augmented with the LLaVA architecture.

    vision 7b

    892.4K  Pulls 17  Tags Updated  2 years ago

  • llava-phi3

    A new small LLaVA model fine-tuned from Phi 3 Mini.

    vision 3.8b

    399.8K  Pulls 4  Tags Updated  2 years ago

  • Me7war/Astria

    Astria is a multimodal model built by combining a LLaVA vision encoder with the new Ministral model, producing a unified system capable of detailed visual understanding and strong general-purpose reasoning.

    vision tools 4b 8b

    500  Pulls 3  Tags Updated  9 months ago

  • ManishThota/llava_next_video

    LLaVA NeXT Video 7B DPO which can process video and multiple images at once

    vision

    3,982  Pulls 1  Tag Updated  2 years ago

  • 0ssamaak0/xtuner-llava

    Family of LLaVA models fine-tuned from Llama3-8B Instruct, Phi3-mini and CLIP-ViT-Large-patch14-336 with ShareGPT4V-PT and InternVL-SFT by XTuner.

    vision

    3,377  Pulls 4  Tags Updated  2 years ago

  • mannix/llava-phi3

    A new small LLaVA model fine-tuned from Phi 3 Mini [I-Quants]

    vision

    2,056  Pulls 4  Tags Updated  2 years ago

  • injoon5/shoe-explainer

    A model that combines existing data with the LLaVA model's results to explain shoes accurately.

    vision

    1,596  Pulls 1  Tag Updated  2 years ago

  • mapler/llava-mistral

    A new LLaVA model fine-tuned from Mistral's 7B model

    701  Pulls 1  Tag Updated  2 years ago

  • aimagelab/llava-more-8b

    LLaVA built with LLaMA 3.1 8B as LLM

    vision

    158  Pulls 1  Tag Updated  1 year ago

  • AgentricAi/AgentricAI_LLaVa

    is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding.

    vision

    148  Pulls 1  Tag Updated  1 year ago

  • nemotron3

    NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.

    vision tools thinking 33b

    665.1K  Pulls 4  Tags Updated  4 months ago

  • JorgeAtLLama/thematrix

    TheMatrix is a language model built for fans who are passionate about the movie The Matrix and the vast universe of science fiction films. Whether you want to explore the philosophical depth of The Matrix.

    tools

    322  Pulls 1  Tag Updated  1 year ago

  • salmatrafi/acegpt

    AceGPT: Aligning Large Language Models with Local (Arabic) Values

    7b 13b

    18.8K  Pulls 2  Tags Updated  2 years ago

  • lazarevtill/Holo1-7B

    Holo1 is an Action Vision-Language Model (VLM) developed by HCompany for use in the Surfer-H web agent system.

    vision

    278  Pulls 3  Tags Updated  1 year ago

  • u1i/LLaMA-3-MERaLiON-8B-Instruct

    Q4 version of the MERaLiON LLM by I2R and A*STAR. See https://www.meralion.ai/

    204  Pulls 1  Tag Updated  1 year ago

  • JorgeAtLLama/panchovilla

    Pancho Villa is a language model designed to help users explore the rich history of Mexico. Named after the famous Mexican revolutionary leader, Pancho Villa provides a detailed and engaging look into the events, cultures, and figures that shaped Mexico.

    tools

    29  Pulls 1  Tag Updated  1 year ago

  • demonbyron/embeddinggemma-300m-lawvault

    EmbeddingGemma-300M-LawVault is a high-performance embedding model fine-tuned specifically for Chinese Legal RAG (Retrieval-Augmented Generation) scenarios.

    embedding

    206  Pulls 1  Tag Updated  9 months ago

  • z-uo/llava-med-v1.5-mistral-7b_q8_0

    This is a 8 bit quantized version of Large Language and Vision Assistant for bio Medicine

    vision tools

    646  Pulls 1  Tag Updated  1 year ago

© 2026 Ollama
Blog Support