Flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL.
5.1M Pulls 17 Tags Updated 1 year ago
The most powerful vision-language model in the Qwen model family to date.
6.2M Pulls 57 Tags Updated 10 months ago
MiniCPM-V surpasses proprietary models such as GPT-4V, Gemini Pro, Qwen-VL and Claude 3 in overall performance, and support multimodal conversation for over 30 languages.
66K Pulls 8 Tags Updated 2 years ago
Qwen/Qwen3-VL-30B-A3B-Instruct - IQ4_NL Quant
135 Pulls 1 Tag Updated 3 weeks ago
1,233 Pulls 14 Tags Updated 1 month ago
153 Pulls 1 Tag Updated 4 months ago
46.3K Pulls 16 Tags Updated 10 months ago
8,001 Pulls 2 Tags Updated 9 months ago
The Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model.
3,701 Pulls 3 Tags Updated 8 months ago
The Qwen3-VL-Embedding model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model.
3,007 Pulls 3 Tags Updated 8 months ago
2,523 Pulls 1 Tag Updated 9 months ago
25 Pulls 1 Tag Updated 1 month ago
German-OCR-Turbo ist ein fine-tuned Vision-Language-Modell basierend auf Qwen3-VL-2B, optimiert für die präzise Texterkennung aus deutschen Rechnungen, Formularen und Geschäftsdokumenten. Das Modell extrahiert strukturierte Daten im Markdown-Format.
1,873 Pulls 1 Tag Updated 9 months ago
1,019 Pulls 1 Tag Updated 11 months ago
Qwen 3 VL but like Heretic.
712 Pulls 4 Tags Updated 7 months ago
A specialized document classification model based on Qwen2.5-VL-3B that automatically detects document types from PDFs and images with high accuracy and calibrated confidence scores.
608 Pulls 2 Tags Updated 8 months ago
Alibaba Tongyi GUI agent on Qwen3-VL. SOTA: 73.5% ScreenSpot-Pro, 76.7% AndroidWorld. Returns bbox [x1,y1,x2,y2] for UI automation. Supports MCP tools & device-cloud collaboration. Apache 2.0. Tags: 2b (default), 8b.
617 Pulls 3 Tags Updated 8 months ago
High Quality Vision Instruct Model
570 Pulls 1 Tag Updated 7 months ago
485 Pulls 2 Tags Updated 7 months ago
Recommended for transcribing and summarizing text from screenshots.
470 Pulls 1 Tag Updated 8 months ago