1 15 hours ago

Baidu Unlimited-OCR (3.3B): document layout OCR that returns every text region with its bounding box.

vision
ollama run richardyoung/unlimited-ocr

Models

View all →

Readme

Unlimited-OCR

Baidu Unlimited-OCR (3.3B): document layout OCR that returns every text region with its bounding box.

🚀 Overview

GGUF build of baidu/Unlimited-OCR, Baidu’s 3.3B document-layout OCR model that returns each text region with its type and bounding box. It ships with its vision projector, so image input works in Ollama.

Output format: each region comes back as type [x1, y1, x2, y2]text, for example title [34, 88, 370, 172]INVOICE 4471. Pictures are marked as image regions rather than described; use a general vision model such as minicpm-v-4.6 for descriptions.

🎯 Key Features

  • Layout OCR: every region with type and bounding box
  • Marks pictures as image regions
  • Structured output for downstream parsing
  • MIT license

🏷️ Available Versions

Tag Size (+ 0.83 GB projector) BPW Notes
latest / Q4_K_M 1.95 GB 4.85 Recommended
Q8_0 3.13 GB 8.5 Near-lossless

💻 Quick Start

ollama run richardyoung/unlimited-ocr
# then drag an image file into the prompt, or send it in the API "images" field

🛠️ Use Cases

  • Document layout analysis
  • OCR with coordinates for highlighting or redaction
  • Preprocessing documents for RAG

✅ Verified

Tested through Ollama before upload with a synthetic invoice and a shapes image: it transcribed the whole invoice with a bounding box per line (e.g. title [34, 88, 370, 172]INVOICE 4471) and marked both shapes as image regions at the right coordinates.

🔧 Technical Details

  • Base Model: baidu/Unlimited-OCR
  • Parameters: 3.3B
  • Context Length: 32K tokens
  • Quantization: GGUF via llama.cpp; vision projector kept at F16
  • License: MIT (inherited from the base model)

🙏 Acknowledgments

  • Base Model: Baidu
  • Quantization: llama.cpp

Built & maintained by Richard Young · DeepNeuro