A general-purpose multimodal mixture-of-experts model for production-grade tasks and enterprise workloads.
105.6K Pulls 1 Tag Updated 9 months ago
Mistral Small 3 sets a new benchmark in the “small” Large Language Models category below 70B.
3.1M Pulls 21 Tags Updated 1 year ago
Mistral Large 2 is Mistral's new flagship model that is significantly more capable in code generation, mathematics, and reasoning with 128k context window and support for dozens of languages.
1.3M Pulls 32 Tags Updated 1 year ago
Mistral Medium 3.5 is the first flagship model of Mistral AI that merged instruction-following, reasoning, and coding in a single set of 128B weights.
382.9K Pulls 5 Tags Updated 4 months ago
Stable Code 3B is a coding model with instruct and code completion variants on par with models such as Code Llama 7B that are 2.5x larger.
1.1M Pulls 36 Tags Updated 2 years ago
41 Pulls 1 Tag Updated 4 months ago
Large dataset roleplaying model finetuned for natural response. Made by ReadyArt (Huggingface).
1,254 Pulls 9 Tags Updated 9 months ago
69.2K Pulls 1 Tag Updated 2 years ago
20.3K Pulls 10 Tags Updated 1 year ago
3 Pulls 1 Tag Updated 1 month ago
Quantized variants of a German large language model (LLM).
2,149 Pulls 12 Tags Updated 2 years ago
EM German is a Llama2/Mistral/LeoLM-based model family, finetuned on a large dataset of various instructions in German language. From https://github.com/jphme/EM_German.
1,898 Pulls 1 Tag Updated 2 years ago
This is a 8 bit quantized version of Large Language and Vision Assistant for bio Medicine
644 Pulls 1 Tag Updated 1 year ago
Mistral Small 3 sets a new benchmark in the “small” Large Language Models category below 70B. This model is unsloth version of the model.
218 Pulls 1 Tag Updated 1 year ago
I-quants for mistral-large-instruct-2407
202 Pulls 7 Tags Updated 1 year ago
Large Language and Vision Assistant for bio Medicine
168 Pulls 1 Tag Updated 1 year ago
IQ2_XSS quant of mistral large 2
173 Pulls 1 Tag Updated 2 years ago
28 Pulls 1 Tag Updated 2 years ago
Large context
8 Pulls 1 Tag Updated 1 year ago
Qwen2.5 models are pretrained on Alibaba's latest large-scale dataset, encompassing up to 18 trillion tokens. The model supports up to 128K tokens and has multilingual capabilities. The following model is specialized on Cline (previously Claude-dev)
1,331 Pulls 1 Tag Updated 1 year ago