65 Downloads Updated 1 week ago
ollama run hisparshmishra1/fast-qwen
Qwen2.5-1.5B-Instruct served fast on modest hardware — a benchmark-driven optimization study, not a retrain.
Same weights as the official Q4_K_M quantization of Qwen2.5-1.5B-Instruct, packaged with a tuned ChatML template and a 4K context default.
| Configuration | Speed | Perplexity (lower = better) |
|---|---|---|
| Original FP16 | 14.4 tok/s | 7.877 |
| Q8_0 (zero measurable loss) | 55.2 tok/s | 7.881 |
| Q4_K_M — this model | ~78 tok/s (5.4×) | 8.114 (+3%) |
Quality verified three ways (identical inputs across all configs):
Q4_K_M quantization → full GPU offload → FlashAttention → n-gram speculative decoding → 8-bit KV cache.
Honest negative result: classic draft-model speculative decoding made this model size slower (27 vs 78 tok/s) — it pays off for large, slow models, not a 1.5B sharing a 4GB GPU with the desktop. Full methodology, raw logs, and a reproducible benchmark harness live in the fast-llm project repository.
ollama run hisparshmishra1/fast-qwen
Tip: close GPU-heavy apps before chatting — a 4GB card shared with the desktop may split the model across CPU/GPU (~40 tok/s) instead of running fully on GPU (~78 tok/s).