DeepSeek-V3.1-Terminus is a hybrid model that supports both thinking mode and non-thinking mode.
726.9K Pulls 7 Tags Updated 11 months ago
A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
3.8M Pulls 5 Tags Updated 1 year ago
Senior Go & SpecKit engineering agent powered by DeepSeek-v3.1 671B, optimized for idiomatic development and deterministic BDD testing.
140 Pulls 1 Tag Updated 5 months ago
38 Pulls 1 Tag Updated 4 months ago
This model is a distilled version of Qwen/Qwen3-30B-A3B-Instruct designed to inherit the reasoning and behavioral characteristics of its much larger teacher model, deepseek-ai/DeepSeek-V3.1.
2,373 Pulls 2 Tags Updated 1 year ago
This is not the ablation version. DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode.
253 Pulls 3 Tags Updated 1 year ago
188 Pulls 2 Tags Updated 11 months ago
15 Pulls 1 Tag Updated 1 year ago
DeepSeek-V3-Pruned-Coder-411B is a pruned version of the DeepSeek-V3 reduced from 256 experts to 160 experts, The pruned model is mainly used for code generation.
1,407 Pulls 5 Tags Updated 1 year ago
67.7K Pulls 17 Tags Updated 3 weeks ago
0731
1,110 Pulls 17 Tags Updated 3 weeks ago
7,196 Pulls 2 Tags Updated 1 year ago
4,108 Pulls 5 Tags Updated 1 year ago
(Unsloth Dynamic Quants) A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
2,064 Pulls 3 Tags Updated 1 year ago
DeepSeep V3 from March 2025 Merged from Unsloth's HF - 671B params - Q8_0/713 GB & Q4_K_M/404 GB
969 Pulls 4 Tags Updated 1 year ago
Single file version with (Dynamic Quants) A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
117 Pulls 4 Tags Updated 1 year ago
25 Pulls 1 Tag Updated 1 year ago
To run DeepSeek-V4-Flash in full precision lossless, run Q3 (UD-Q3_K_M), It is 129 GB. vision tools thinking
144 Pulls 1 Tag Updated 1 month ago