74 1 year ago

A tiny fine-tuned and distillate model from Llama with GRPO for enhance reasoning to use like a personal assistant on smaller devices