825 1 month ago

A memory-efficient model configuration of Qwen3.6-35B-A3B using an upstream imatrix-calibrated IQ4_XS quantization and q4_0 KV cache. Designed for 24 GB VRAM

tools thinking
c8472cd9daed · 31B
You are a helpful AI assistant.