65 1 week ago

Qwen2.5-1.5B-Instruct (Q4_K_M), tuned for fast local inference — 5.4x faster than the original with verified quality (perplexity +3%, task scores unchanged).

ollama run hisparshmishra1/fast-qwen

Models

View all →

1 model

fast-qwen:latest

1.1GB · 32K context window · Text · 1 week ago

Readme

fast-qwen ⚡

Qwen2.5-1.5B-Instruct served fast on modest hardware — a benchmark-driven optimization study, not a retrain.

Same weights as the official Q4_K_M quantization of Qwen2.5-1.5B-Instruct, packaged with a tuned ChatML template and a 4K context default.

Measured results (GTX 1650 4GB, i7-9750HF, llama.cpp)

Configuration Speed Perplexity (lower = better)
Original FP16 14.4 tok/s 7.877
Q8_0 (zero measurable loss) 55.2 tok/s 7.881
Q4_K_M — this model ~78 tok/s (5.4×) 8.114 (+3%)

Quality verified three ways (identical inputs across all configs):

  • Perplexity on a held-out corpus: +3% vs the original FP16 weights
  • 24-question task suite (math, logic, code, facts): 15⁄24 optimized vs 13⁄24 original — no drop
  • Side-by-side outputs: wording drifts slightly, meaning and correctness preserved on every identity-check question

Techniques applied

Q4_K_M quantization → full GPU offload → FlashAttention → n-gram speculative decoding → 8-bit KV cache.

Honest negative result: classic draft-model speculative decoding made this model size slower (27 vs 78 tok/s) — it pays off for large, slow models, not a 1.5B sharing a 4GB GPU with the desktop. Full methodology, raw logs, and a reproducible benchmark harness live in the fast-llm project repository.

Usage

ollama run hisparshmishra1/fast-qwen

Tip: close GPU-heavy apps before chatting — a 4GB card shared with the desktop may split the model across CPU/GPU (~40 tok/s) instead of running fully on GPU (~78 tok/s).