Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
940.4K Pulls 12 Tags Updated 1 week ago
Z.ai's first natively multimodal model, approaching Claude Opus 4.8 on coding and agentic benchmarks with just 18B active parameters.
9,023 Pulls 1 Tag Updated yesterday
DeepSeek-V4-Flash is a preview of the DeepSeek-V4 series, a Mixture-of-Experts model with 284B total parameters and 13B activated, built for efficient reasoning across a 1M-token context window.
400.8K Pulls 2 Tags Updated 3 weeks ago
This experimental preview of the architecture that will underpin Qwen4.
2,211 Pulls 3 Tags Updated 16 hours ago
Meta's latest open model built for always-on local agents. 30B parameters, licensed under Apache 2.0 and runs on a single GPU — tuned for tool use, long tasks, and failure recovery.
170.4K Pulls 15 Tags Updated 1 week ago
NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for always-on agents.
136.6K Pulls 11 Tags Updated 2 weeks ago
Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
47.5K Pulls 3 Tags Updated 1 week ago
IBM Granite Models are a family of enterprise-ready, open foundation models that support multilingual capabilities, coding, retrieval-augmented generation (RAG), tool use, thinking and structured JSON output. Released under Apache 2.0 license.
8,606 Pulls 49 Tags Updated 2 days ago
Gemma 4 models are designed to deliver frontier-level performance at each size. They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.
23.5M Pulls 50 Tags Updated 21 hours ago
Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
6.2M Pulls 35 Tags Updated yesterday
GLM-5.1 is our next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin.
2.3M Pulls 1 Tag Updated 4 months ago
MiniMax's M2-series model for coding, agentic workflows, and professional productivity.
2.4M Pulls 1 Tag Updated 5 months ago
A self-improving family of open-source models for agentic coding
423.4K Pulls 9 Tags Updated 2 months ago
MiniMax M3: Coding & Agentic Frontier. 1M context window. Native Multimodality.
489.2K Pulls 1 Tag Updated 2 months ago
GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks.
332.6K Pulls 1 Tag Updated 2 months ago
NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.
650.6K Pulls 4 Tags Updated 4 months ago
Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built upon Kimi K2.6, with substantial improvements on real-world long-horizon coding tasks and roughly 30% lower thinking-token usage.
225.8K Pulls 1 Tag Updated 2 months ago
IBM Granite Models are a family of enterprise-ready, open foundation models that support multilingual capabilities, coding, retrieval-augmented generation (RAG), tool use, and structured JSON output. Released under Apache 2.0 license.
380K Pulls 48 Tags Updated 3 months ago
Mistral Medium 3.5 is the first flagship model of Mistral AI that merged instruction-following, reasoning, and coding in a single set of 128B weights.
358.1K Pulls 5 Tags Updated 3 months ago
DeepSeek-V4-Pro is a frontier Mixture-of-Experts model with a large context window and three reasoning modes.
360.5K Pulls 2 Tags Updated 2 weeks ago