smtek/ Qwen3.8-27B:IQ3_XXS

1,022 10 hours ago

Quants from Q2 up to Q5 from Unsloth K_M and UD_K_XL

ollama run smtek/Qwen3.8-27B:IQ3_XXS

Details

20 hours ago

96b6078db79a · 12GB ·

qwen35
·
27.3B
·
Q4_K_S
{{ .Prompt }}
{ "num_ctx": 262144, "stop": [ "<|im_start|>", "<|im_end|>" ] }

Readme

Models

Tag Quant Size Context VRAM target Use case
latest 4-bit XL 17.9 GB 128K 24 GB Default (Q4_K_XL)
IQ2_XXS 2-bit 9.0 GB 256K 16 GB Smallest, max context
IQ2_M 2-bit 10 GB 128K 16 GB Tight VRAM, long context
IQ2_M-256k 2-bit 10 GB 256K 24 GB Tight VRAM, native context
IQ2_M-12gb 2-bit 10.3 GB 32K 12 GB 2-bit KM, optimized for 12 GB VRAM
Q2_K_XL-12gb 2-bit 10.7 GB 32K 12 GB 2-bit XL, optimized for 12 GB VRAM
Q2_K_XL-16gb 2-bit 10.7 GB 48K 16 GB 2-bit XL, optimized for 16 GB VRAM
Q2_K_XL 2-bit 10.7 GB 256K 24 GB 2-bit XL, max context
IQ3_XXS 3-bit 11.9 GB 256K 24 GB 3-bit, native context
Q3_K_M 3-bit 13 GB 128K 16 GB Balanced size / quality
Q3_K_M-192k 3-bit 13 GB 192K 20 GB Balanced, longer context
Q3_K_XL 3-bit 13.4 GB 192K 24 GB 3-bit XL, high quality
Q3_K_XL-16gb 3-bit 13.4 GB 48K 16 GB Optimized for 16 GB VRAM
Q4_K_M 4-bit 17 GB 128K 24 GB Best 4-bit quality
Q4_K_XL 4-bit 17.9 GB 128K 24 GB Recommended default for 24 GB
Q5_K_M 5-bit 19 GB 128K 24 GB Highest quality that fits 24 GB
Q5_K_XL 5-bit 20.2 GB 64K 24 GB 5-bit XL, best quality in 24 GB
Q5_K_M-256k 5-bit 19 GB 256K 32 GB Native context, needs 32 GB

All untagged models use num_ctx 131072 and the Qwen im_start/im_end stop tokens. The -256k/-192k/-92k/-64k suffixes set the context window explicitly. The -12gb/-16gb suffixes are context-optimized for those VRAM budgets. The UD- (Unsloth’s MTP implementation) quants use the XL block size for higher quality at the same bit depth.

Usage

ollama run smtek/Qwen3.8-27B:Q4_K_M

Or via the API:

curl http://localhost:11434/api/generate -d '{
  "model": "smtek/Qwen3.8-27B:Q4_K_M",
  "prompt": "Hello"
}'

Recommended Optimization

By default, Ollama allocates the KV-Cache in f16, which will exceed VRAM over long contexts. To run stably up to the full 256K window, set these environment variables in your system (/etc/systemd/system/ollama.service.d/override.conf on Linux):

[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
  • Flash Attention: Dramatically reduces context memory footprint.
  • q8_0 Cache: Compresses context tensors without the severe logic degradation caused by 4-bit (q4_0) cache quantization.

After editing, reload and restart:

sudo systemctl daemon-reload
sudo systemctl restart ollama

Source

  • Hugging Face: unsloth/Qwen3.8-27B-GGUF
  • License: Apache-2.0