MiniMax's M2-series model for coding, agentic workflows, and professional productivity.
2.4M Pulls 1 Tag Updated 5 months ago
Mistral Medium 3.5 is the first flagship model of Mistral AI that merged instruction-following, reasoning, and coding in a single set of 128B weights.
327.3K Pulls 5 Tags Updated 3 months ago
A general-purpose multimodal mixture-of-experts model for production-grade tasks and enterprise workloads.
90.9K Pulls 1 Tag Updated 8 months ago
The 7B model released by Mistral AI, updated to version 0.3.
31.9M Pulls 84 Tags Updated 1 year ago
A state-of-the-art 12B model with 128k context length, built by Mistral AI in collaboration with NVIDIA.
5.6M Pulls 17 Tags Updated 1 year ago
A set of Mixture of Experts (MoE) model with open weights by Mistral AI in 8x7b and 8x22b parameter sizes.
2.8M Pulls 70 Tags Updated 1 year ago
Mistral Large 2 is Mistral's new flagship model that is significantly more capable in code generation, mathematics, and reasoning with 128k context window and support for dozens of languages.
1.3M Pulls 32 Tags Updated 1 year ago
Mistral OpenOrca is a 7 billion parameter model, fine-tuned on top of the Mistral 7B model using the OpenOrca dataset.
671.8K Pulls 17 Tags Updated 2 years ago
MistralLite is a fine-tuned model based on Mistral with enhanced capabilities of processing long contexts.
516.2K Pulls 17 Tags Updated 2 years ago
Mistral Small 3 sets a new benchmark in the “small” Large Language Models category below 70B.
3.1M Pulls 21 Tags Updated 1 year ago
DeepSeek-V4-Flash is a preview of the DeepSeek-V4 series, a Mixture-of-Experts model with 284B total parameters and 13B activated, built for efficient reasoning across a 1M-token context window.
366.9K Pulls 3 Tags Updated 2 weeks ago
DeepSeek-V4-Pro is a frontier Mixture-of-Experts model with a large context window and three reasoning modes.
329.6K Pulls 3 Tags Updated 2 days ago
Laguna XS 2.1 is a 33B total parameter Mixture-of-Experts model with 3B activated parameters per token designed for agentic coding and long-horizon work on a local machine.
90.9K Pulls 7 Tags Updated 3 weeks ago
North Mini Code is Cohere's first model for developers — a 30B Mixture-of-Experts model with 3B active parameters, built for agentic software engineering.
41K Pulls 7 Tags Updated 1 month ago
Laguna XS.2 is a 33B total parameter Mixture-of-Experts model with 3B activated parameters per token designed for agentic coding and long-horizon work on a local machine.
25.4K Pulls 7 Tags Updated 3 weeks ago
Phi-3 is a family of lightweight 3B (Mini) and 14B (Medium) state-of-the-art open models by Microsoft.
18M Pulls 72 Tags Updated 2 years ago
A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
3.8M Pulls 5 Tags Updated 1 year ago
An open-source Mixture-of-Experts code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks.
3M Pulls 64 Tags Updated 1 year ago
Uncensored, 8x7b and 8x22b fine-tuned models based on the Mixtral mixture of experts models that excels at coding tasks. Created by Eric Hartford.
1.9M Pulls 70 Tags Updated 1 year ago
Codestral is Mistral AI’s first-ever code model designed for code generation tasks.
1.3M Pulls 17 Tags Updated 1 year ago