Updated 17 hours ago
ollama run richardyoung/minicpm-v-4.6:Q4_K_M
MiniCPM-V 4.6 vision model (1.3B): image understanding and OCR in about 1.6 GB. Image input works.
GGUF build of openbmb/MiniCPM-V-4.6, OpenBMB’s compact vision-language model: describes images, answers questions about them and reads text in photos and documents. It ships with its vision projector, so image input works in Ollama.
| Tag | Size (+ 1.11 GB projector) | BPW | Notes |
|---|---|---|---|
| latest / Q4_K_M | 0.53 GB | 4.85 | Recommended |
| Q8_0 | 0.81 GB | 8.5 | Near-lossless |
ollama run richardyoung/minicpm-v-4.6
# then drag an image file into the prompt, or send it in the API "images" field
Tested through Ollama before upload with a synthetic invoice and a shapes image: it transcribed the whole invoice exactly and named both shapes and their colors.
--no-mtp (the release declares a multi-token-prediction layer it does not ship)Built & maintained by Richard Young · DeepNeuro