74 1 year ago

A tiny fine-tuned and distillate model from Llama with GRPO for enhance reasoning to use like a personal assistant on smaller devices

fa956ab37b8c · 98B
{
"stop": [
"<|system|>",
"<|user|>",
"<|assistant|>",
"</s>"
]
}