24 Downloads Updated yesterday
ollama run bluehawana/deepseek-v4-flash:iq2_m
Updated yesterday
yesterday
53fa62a979e3 · 104GB ·
Single-file IQ2_M quant of DeepSeek-V4-Flash-0731 for Ollama. 284B mixture-of-experts, 13B active per token, 1M context, MIT license.
Benchmarked on a MacBook Pro M5 Max (128 GB unified memory): ~33 tok/s generation, ~545 tok/s prompt processing, 3.9 s warm load.
Tip: 104 GB of weights in 128 GB RAM leaves little KV-cache headroom. Keep context modest: /set parameter num_ctx 8192
Credits: this is not our model. Base model by DeepSeek (MIT), quantization and imatrix by AtomicChat (huggingface.co/AtomicChat/DeepSeek-V4-Flash-0731-GGUF). We merged their four GGUF shards into one file (llama-gguf-split –merge, byte-preserving) and re-hosted it so Ollama can pull it.
Scripts, the do-it-yourself recipe for any sharded GGUF, and Mac tips: github.com/bluehawana/ollama-deepseek-v4-flash-iq2_m-mac-m5max
Also on Hugging Face: huggingface.co/bluehawana/DeepSeek-V4-Flash-0731-IQ2_M-GGUF