17 hours ago

MiniCPM-V 4.6 vision model (1.3B): image understanding and OCR in about 1.6 GB. Image input works.

vision
ollama run richardyoung/minicpm-v-4.6:Q4_K_M

Details

17 hours ago

c5d827b4c7e0 · 1.6GB

qwen35
·
752M
·
Q4_K_M
clip
·
548M
·
F16

Readme

MiniCPM-V-4.6

MiniCPM-V 4.6 vision model (1.3B): image understanding and OCR in about 1.6 GB. Image input works.

🚀 Overview

GGUF build of openbmb/MiniCPM-V-4.6, OpenBMB’s compact vision-language model: describes images, answers questions about them and reads text in photos and documents. It ships with its vision projector, so image input works in Ollama.

🎯 Key Features

  • Image understanding and visual Q&A
  • OCR of photos, screenshots and documents
  • Tiny: about 1.6 GB with the vision projector
  • Apache-2.0 license

🏷️ Available Versions

Tag Size (+ 1.11 GB projector) BPW Notes
latest / Q4_K_M 0.53 GB 4.85 Recommended
Q8_0 0.81 GB 8.5 Near-lossless

💻 Quick Start

ollama run richardyoung/minicpm-v-4.6
# then drag an image file into the prompt, or send it in the API "images" field

🛠️ Use Cases

  • Local image description and visual Q&A
  • Screenshot and document reading
  • Low-VRAM and CPU vision

✅ Verified

Tested through Ollama before upload with a synthetic invoice and a shapes image: it transcribed the whole invoice exactly and named both shapes and their colors.

🔧 Technical Details

  • Base Model: openbmb/MiniCPM-V-4.6
  • Parameters: 1.3B
  • Context Length: 256K tokens
  • Quantization: GGUF via llama.cpp; vision projector kept at F16; converted with --no-mtp (the release declares a multi-token-prediction layer it does not ship)
  • License: Apache-2.0 (inherited from the base model)

🙏 Acknowledgments

  • Base Model: OpenBMB
  • Quantization: llama.cpp

Built & maintained by Richard Young · DeepNeuro