168 1 week ago

Qwen 3.6 35B multimodal model with vision support, quantized to IQ3_S and optimized to fit within 16 GB VRAM with Q4_0 KV cache. Supports very large context windows, with the practical maximum depending on the Ollama version, GPU and backend.

vision tools thinking 35b
86eff881e8d2 · 94B
{
"min_p": 0,
"presence_penalty": 1.5,
"repeat_penalty": 1,
"temperature": 1,
"top_k": 20,
"top_p": 0.95
}