A fine-tuned version of Deepseek-R1-Distilled-Qwen-1.5B that surpasses the performance of OpenAI’s o1-preview with just 1.5B parameters on popular math evaluations.
1.3M Pulls 5 Tags Updated 1 year ago
1,967 Pulls 3 Tags Updated 1 year ago
226 Pulls 1 Tag Updated 1 year ago
DeepScaleR-1.5B-Preview is a language model fine-tuned from DeepSeek-R1-Distilled-Qwen-1.5B using distributed reinforcement learning (RL)
110 Pulls 1 Tag Updated 1 year ago
6 Pulls 1 Tag Updated 1 year ago