1 Download Updated 15 hours ago
ollama run richardyoung/unlimited-ocr
Baidu Unlimited-OCR (3.3B): document layout OCR that returns every text region with its bounding box.
GGUF build of baidu/Unlimited-OCR, Baidu’s 3.3B document-layout OCR model that returns each text region with its type and bounding box. It ships with its vision projector, so image input works in Ollama.
Output format: each region comes back as type [x1, y1, x2, y2]text, for example title [34, 88, 370, 172]INVOICE 4471. Pictures are marked as image regions rather than described; use a general vision model such as minicpm-v-4.6 for descriptions.
| Tag | Size (+ 0.83 GB projector) | BPW | Notes |
|---|---|---|---|
| latest / Q4_K_M | 1.95 GB | 4.85 | Recommended |
| Q8_0 | 3.13 GB | 8.5 | Near-lossless |
ollama run richardyoung/unlimited-ocr
# then drag an image file into the prompt, or send it in the API "images" field
Tested through Ollama before upload with a synthetic invoice and a shapes image: it transcribed the whole invoice with a bounding box per line (e.g. title [34, 88, 370, 172]INVOICE 4471) and marked both shapes as image regions at the right coordinates.
Built & maintained by Richard Young · DeepNeuro