99 1 week ago

Qwen 3.6 35B text-only model, quantized to IQ3_S and optimized to fit within 16 GB VRAM with Q4_0 KV cache. Designed for long-context chat, coding, document analysis and general-purpose local inference. Supports very large context windows.

tools thinking 35b
86eff881e8d2 · 94B
{
"min_p": 0,
"presence_penalty": 1.5,
"repeat_penalty": 1,
"temperature": 1,
"top_k": 20,
"top_p": 0.95
}