19 Downloads Updated yesterday
ollama run MobiusDevelopment/Qwen3.6-35B-A3B-A8E4-128K
Mixed-precision quant of Qwen/Qwen3.6-35B-A3B: 8-bit attention, 4-bit experts. Tuned to run entirely on a single 24 GB GPU with a 128K context.
Attention runs on every token, so its quantization error adds up. Each expert only sees a few tokens, so experts tolerate 4-bit well (see MoQE). A8E4 keeps attention, Gated DeltaNet, shared experts and embeddings at Q8_0, and quantizes only the routed experts (ffn_{gate,up,down}_exps) to Q4_K.
The result is smaller than a standard Q4_K_M (19.4 vs 19.7 GiB) and much closer to Q8_0.
| vs Q8_0, code corpus | A8E4 | Q4_K_M |
|---|---|---|
| Mean KL divergence | 0.034 | 0.087 |
| 99th-percentile KL divergence | 0.43 | 1.05 |
| Perplexity change | +0.9% | +2.6% |
| Same top token | 96.3% | 94.7% |
RTX 3090 Ti 24 GB, Ollama 0.35.0, flash attention on, q8_0 KV cache, 128K context:
num_gpu 999 and num_ctx 131072 so it always loads fully on the GPU. On smaller cards it may fail to load: lower num_ctx or remove num_gpu.OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
ollama run MobiusDevelopment/Qwen3.6-35B-A3B-A8E4-128K
OpenAI-compatible endpoint: http://localhost:11434/v1, model MobiusDevelopment/Qwen3.6-35B-A3B-A8E4-128K.
Default sampling: temperature 0.6, top_p 0.95, top_k 20, min_p 0.
llama-quantize --tensor-type, no imatrix.Apache 2.0, inherited from Qwen3.6-35B-A3B.