🌋 LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.
15M Pulls 98 Tags Updated 2 years ago
A LLaVA model fine-tuned from Llama 3 Instruct with better scores in several benchmarks.
2.3M Pulls 4 Tags Updated 2 years ago
BakLLaVA is a multimodal model consisting of the Mistral 7B base model augmented with the LLaVA architecture.
892.4K Pulls 17 Tags Updated 2 years ago
A new small LLaVA model fine-tuned from Phi 3 Mini.
399.8K Pulls 4 Tags Updated 2 years ago
Astria is a multimodal model built by combining a LLaVA vision encoder with the new Ministral model, producing a unified system capable of detailed visual understanding and strong general-purpose reasoning.
500 Pulls 3 Tags Updated 9 months ago
LLaVA NeXT Video 7B DPO which can process video and multiple images at once
3,982 Pulls 1 Tag Updated 2 years ago
Family of LLaVA models fine-tuned from Llama3-8B Instruct, Phi3-mini and CLIP-ViT-Large-patch14-336 with ShareGPT4V-PT and InternVL-SFT by XTuner.
3,377 Pulls 4 Tags Updated 2 years ago
A new small LLaVA model fine-tuned from Phi 3 Mini [I-Quants]
2,056 Pulls 4 Tags Updated 2 years ago
A model that combines existing data with the LLaVA model's results to explain shoes accurately.
1,596 Pulls 1 Tag Updated 2 years ago
A new LLaVA model fine-tuned from Mistral's 7B model
701 Pulls 1 Tag Updated 2 years ago
LLaVA built with LLaMA 3.1 8B as LLM
158 Pulls 1 Tag Updated 1 year ago
is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding.
148 Pulls 1 Tag Updated 1 year ago
NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.
665.1K Pulls 4 Tags Updated 4 months ago
TheMatrix is a language model built for fans who are passionate about the movie The Matrix and the vast universe of science fiction films. Whether you want to explore the philosophical depth of The Matrix.
322 Pulls 1 Tag Updated 1 year ago
AceGPT: Aligning Large Language Models with Local (Arabic) Values
18.8K Pulls 2 Tags Updated 2 years ago
Holo1 is an Action Vision-Language Model (VLM) developed by HCompany for use in the Surfer-H web agent system.
278 Pulls 3 Tags Updated 1 year ago
Q4 version of the MERaLiON LLM by I2R and A*STAR. See https://www.meralion.ai/
204 Pulls 1 Tag Updated 1 year ago
Pancho Villa is a language model designed to help users explore the rich history of Mexico. Named after the famous Mexican revolutionary leader, Pancho Villa provides a detailed and engaging look into the events, cultures, and figures that shaped Mexico.
29 Pulls 1 Tag Updated 1 year ago
EmbeddingGemma-300M-LawVault is a high-performance embedding model fine-tuned specifically for Chinese Legal RAG (Retrieval-Augmented Generation) scenarios.
206 Pulls 1 Tag Updated 9 months ago
This is a 8 bit quantized version of Large Language and Vision Assistant for bio Medicine
646 Pulls 1 Tag Updated 1 year ago