Alibaba's performant long context models for agentic and coding tasks.
9.3M Pulls 9 Tags Updated 11 months ago
The latest series of Code-Specific Qwen models, with significant improvements in code generation, code reasoning, and code fixing.
21.5M Pulls 199 Tags Updated 1 year ago
Qwen3-Coder-Next is a coding-focused language model from Alibaba's Qwen team, optimized for agentic coding workflows and local development.
2M Pulls 3 Tags Updated 7 months ago
111 billion parameter model optimized for demanding enterprises that require fast, secure, and high-quality AI
229.1K Pulls 5 Tags Updated 1 year ago
Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built upon Kimi K2.6, with substantial improvements on real-world long-horizon coding tasks and roughly 30% lower thinking-token usage.
239.1K Pulls 1 Tag Updated 3 months ago
Mistral Large 2 is Mistral's new flagship model that is significantly more capable in code generation, mathematics, and reasoning with 128k context window and support for dozens of languages.
1.3M Pulls 32 Tags Updated 1 year ago
Cogito v1 Preview is a family of hybrid reasoning models by Deep Cogito that outperform the best available open models of the same size, including counterparts from LLaMA, DeepSeek, and Qwen across most standard benchmarks.
2.1M Pulls 20 Tags Updated 1 year ago
Command R+ is a powerful, scalable large language model purpose-built to excel at real-world enterprise use cases.
804.2K Pulls 21 Tags Updated 2 years ago
The smallest model in Cohere's R series delivers top-tier speed, efficiency, and quality to build powerful AI applications on commodity GPUs and edge devices.
323.9K Pulls 5 Tags Updated 1 year ago
Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
6.7M Pulls 35 Tags Updated 2 weeks ago
MiniMax M3: Coding & Agentic Frontier. 1M context window. Native Multimodality.
511.2K Pulls 1 Tag Updated 3 months ago
North Mini Code is Cohere's first model for developers — a 30B Mixture-of-Experts model with 3B active parameters, built for agentic software engineering.
55.8K Pulls 7 Tags Updated 3 months ago
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.
7.3M Pulls 3 Tags Updated 7 months ago
24B model that excels at using tools to explore codebases, editing multiple files and power software engineering agents.
1.1M Pulls 6 Tags Updated 9 months ago
Rnj-1 is a family of 8B parameter open-weight, dense models trained from scratch by Essential AI, optimized for code and STEM with capabilities on par with SOTA open-weight models.
508.9K Pulls 5 Tags Updated 9 months ago
123B model that excels at using tools to explore codebases, editing multiple files and power software engineering agents.
356.3K Pulls 5 Tags Updated 9 months ago
A compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more.
1M Pulls 5 Tags Updated 1 year ago
Athene-V2 is a 72B parameter model which excels at code completion, mathematics, and log extraction tasks.
593.1K Pulls 17 Tags Updated 1 year ago
An open 30B MoE model from NVIDIA with 3B activated parameters that delivers strong reasoning and agentic capabilities.
148K Pulls 3 Tags Updated 6 months ago
Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
2.2M Pulls 12 Tags Updated 1 month ago