DeepSeek-V4.1-Flash is an advanced tool designed to enhance search capabilities, providing users with faster and more accurate results.
39.8K Pulls 1 Tag Updated 1 week ago
This experimental preview of the architecture that will underpin Qwen4.
109.9K Pulls 6 Tags Updated 2 weeks ago
Z.ai's first natively multimodal model, approaching Claude Opus 4.8 on coding and agentic benchmarks with just 18B active parameters.
145.2K Pulls 1 Tag Updated 3 weeks ago
Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
329.1K Pulls 3 Tags Updated 1 month ago
Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
2.3M Pulls 12 Tags Updated 1 month ago
Meta's latest open model built for always-on local agents. 30B parameters, licensed under Apache 2.0 and runs on a single GPU — tuned for tool use, long tasks, and failure recovery.
219.2K Pulls 15 Tags Updated 3 weeks ago
Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date.
85.7K Pulls 1 Tag Updated 1 month ago
Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built upon Kimi K2.6, with substantial improvements on real-world long-horizon coding tasks and roughly 30% lower thinking-token usage.
240.2K Pulls 1 Tag Updated 3 months ago
MiniMax M3: Coding & Agentic Frontier. 1M context window. Native Multimodality.
512.3K Pulls 1 Tag Updated 3 months ago
A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone
40.7K Pulls 13 Tags Updated 3 months ago
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
51.6K Pulls 13 Tags Updated 3 months ago
Mistral Medium 3.5 is the first flagship model of Mistral AI that merged instruction-following, reasoning, and coding in a single set of 128B weights.
404.8K Pulls 5 Tags Updated 4 months ago
NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.
663.5K Pulls 4 Tags Updated 4 months ago
Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
6.7M Pulls 35 Tags Updated 2 weeks ago
MedGemma 1.5 4B is an updated version of the MedGemma 4B model.
189.7K Pulls 5 Tags Updated 5 months ago
MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension.
402.1K Pulls 9 Tags Updated 5 months ago
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
476.2K Pulls 1 Tag Updated 5 months ago
Gemma 4 models are designed to deliver frontier-level performance at each size. They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.
25.4M Pulls 50 Tags Updated 2 weeks ago
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.
20.6M Pulls 64 Tags Updated 2 weeks ago
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.
7.4M Pulls 3 Tags Updated 7 months ago