24 Pulls 1 Tag Updated 1 year ago
19 Pulls 1 Tag Updated 3 days ago
DeepSeek-R1-0528-Qwen3-8B
1,145 Pulls 1 Tag Updated 3 months ago
DeepSeek-R1-0528-Qwen3-8B-IQ4_NL
3,884 Pulls 1 Tag Updated 1 year ago
DeepSeek R1 0528 Qwen3 8B with tool calling/MCP support
3,209 Pulls 1 Tag Updated 1 year ago
DeepSeek R1 Distilled model to one-fourth its original file size—without losing any accuracy.
3,458 Pulls 1 Tag Updated 1 year ago
DeepSeek R1 0528 Qwen3 8B Q4 with tool calling
2,349 Pulls 1 Tag Updated 1 year ago
DeepSeek-R1-0528-Qwen3-8B,包含2个量化模型:Q5_K_M,Q8_0
1,762 Pulls 2 Tags Updated 1 year ago
1,675 Pulls 1 Tag Updated 1 year ago
unsloth微调DeepSeek-R1-Distill-Llama-8B
174 Pulls 1 Tag Updated 1 year ago
Trained on Dataset: https://huggingface.co/datasets/lumolabs-ai/Lumo-Iris-DS-Instruct
112 Pulls 1 Tag Updated 1 year ago
108 Pulls 1 Tag Updated 1 year ago
DeepSeek-R1-0528 仍然使用 2024 年 12 月所发布的 DeepSeek V3 Base 模型作为基座,但在后训练过程中投入了更多算力,显著提升了模型的思维深度与推理能力。这个8B精馏版本编程能力都爆表!
1,086 Pulls 1 Tag Updated 1 year ago
287 Pulls 1 Tag Updated 1 year ago
Unsloth's DeepSeek-R1 1.58-bit, I just merged the thing and uploaded it here. This is the full 671b model, albeit dynamically quantized to 1.58bits.
101.6K Pulls 1 Tag Updated 1 year ago
237 Pulls 1 Tag Updated 1 year ago
68 Pulls 1 Tag Updated 1 year ago
29 Pulls 1 Tag Updated 1 year ago
12 Pulls 1 Tag Updated 1 year ago
5 Pulls 1 Tag Updated 1 year ago