65 1 week ago

Qwen2.5-1.5B-Instruct (Q4_K_M), tuned for fast local inference — 5.4x faster than the original with verified quality (perplexity +3%, task scores unchanged).

e1d53abb2a83 · 92B
{
"num_ctx": 4096,
"stop": [
"<|im_start|>",
"<|im_end|>"
],
"temperature": 0.7
}