Reasoning model based on Phi-4, wich is a 14B parameter, state-of-the-art open model from Microsoft and Unsloth-prepared version of Phi-4 for GRPO & RL process.
63 Pulls 1 Tag Updated 1 year ago