-
deepseek-r1
DeepSeek-R1 is a family of open reasoning models with performance approaching that of leading models, such as O3 and Gemini 2.5 Pro.
tools thinking 1.5b 7b 8b 14b 32b 70b 671b91.6M Pulls 35 Tags Updated 1 year ago
-
deepseek-coder
DeepSeek Coder is a capable coding model trained on two trillion code and natural language tokens.
1.3b 6.7b 33b4.5M Pulls 102 Tags Updated 2 years ago
-
deepseek-v3
A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
671b3.8M Pulls 5 Tags Updated 1 year ago
-
deepseek-coder-v2
An open-source Mixture-of-Experts code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks.
16b 236b3M Pulls 64 Tags Updated 1 year ago
-
deepseek-llm
An advanced language model crafted with 2 trillion bilingual tokens.
7b 67b1.2M Pulls 64 Tags Updated 2 years ago
-
deepseek-v2
A strong, economical, and efficient Mixture-of-Experts language model.
16b 236b1.2M Pulls 34 Tags Updated 2 years ago
-
deepseek-v3.1
DeepSeek-V3.1-Terminus is a hybrid model that supports both thinking mode and non-thinking mode.
tools thinking 671b722K Pulls 7 Tags Updated 10 months ago
-
deepseek-ocr
DeepSeek-OCR is a vision-language model that can perform token-efficient OCR.
vision 3b516.4K Pulls 3 Tags Updated 9 months ago
-
deepseek-v4-flash
DeepSeek-V4-Flash is a preview of the DeepSeek-V4 series, a Mixture-of-Experts model with 284B total parameters and 13B activated, built for efficient reasoning across a 1M-token context window.
tools thinking cloud381.4K Pulls 3 Tags Updated 2 weeks ago
-
deepseek-v4-pro
DeepSeek-V4-Pro is a frontier Mixture-of-Experts model with a large context window and three reasoning modes.
tools thinking cloud345.2K Pulls 3 Tags Updated 6 days ago
-
deepseek-v2.5
An upgraded version of DeekSeek-V2 that integrates the general and coding abilities of both DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct.
236b285.6K Pulls 7 Tags Updated 1 year ago