Qwen2.5-1.5B-Instruct (Q4_K_M), tuned for fast local inference — 5.4x faster than the original with verified quality (perplexity +3%, task scores unchanged).
65 Pulls 1 Tag Updated 1 week ago