mannix/ omnimerge-v4-mtp:vision-Q4_K_L

1,086 1 week ago

Qwen/Qwen3.6-27B + 3 Qwen3.6 fine-tunes with MLP-passthrough surgery - MTP quants

vision tools thinking
ollama run mannix/omnimerge-v4-mtp:vision-Q4_K_L

Details

1 week ago

0708ed4ee648 · 19GB ·

qwen35
·
27.3B
·
Q4_K_M
clip
·
461M
·
F16
{ "draft_num_predict": 3, "min_p": 0, "presence_penalty": 0, "repeat_penalty": 1,

Readme

GGUF quantizations of ManniX-ITA/Qwen3.6-27B-Omnimerge-v4 with the MTP (Multi-Token Prediction) head retained for self-speculative decoding on llama.cpp mainline (PR #22673, merged 2026-05-16) and later.

Up to 4x inference speed with 1 session, 2x with 2 in parallel.

image.png