Japanese instruction-tuned LLM by CyberAgent, distilled from Qwen-72B.
306 Pulls 1 Tag Updated 1 year ago
Tiny-R1-32B-Preview, which outperforms the 70B model Deepseek-R1-Distill-Llama-70B and nearly matches the full R1 model in math.
1,705 Pulls 6 Tags Updated 1 year ago
UNOFFICIAL uploads of the DeepSeek Math 7B RL models
4,990 Pulls 5 Tags Updated 2 years ago