19 yesterday

Mixed-precision quant of Qwen3.6-35B-A3B: 8-bit attention, 4-bit experts.

ollama run MobiusDevelopment/Qwen3.6-35B-A3B-A8E4-128K

Models

View all →

Readme

Qwen3.6-35B-A3B A8E4 (128K)

Mixed-precision quant of Qwen/Qwen3.6-35B-A3B: 8-bit attention, 4-bit experts. Tuned to run entirely on a single 24 GB GPU with a 128K context.

Why A8E4

Attention runs on every token, so its quantization error adds up. Each expert only sees a few tokens, so experts tolerate 4-bit well (see MoQE). A8E4 keeps attention, Gated DeltaNet, shared experts and embeddings at Q8_0, and quantizes only the routed experts (ffn_{gate,up,down}_exps) to Q4_K.

The result is smaller than a standard Q4_K_M (19.4 vs 19.7 GiB) and much closer to Q8_0.

vs Q8_0, code corpus A8E4 Q4_K_M
Mean KL divergence 0.034 0.087
99th-percentile KL divergence 0.43 1.05
Perplexity change +0.9% +2.6%
Same top token 96.3% 94.7%

Measured speed

RTX 3090 Ti 24 GB, Ollama 0.35.0, flash attention on, q8_0 KV cache, 128K context:

  • Generation: ~100 tok/s (87 tok/s at 64K depth, measured with llama.cpp)
  • Prompt processing: ~4,300 tok/s (4096-token prompt, llama.cpp)
  • VRAM: 22.7 GB, 100% GPU

Requirements

  • 24 GB of VRAM. The model sets num_gpu 999 and num_ctx 131072 so it always loads fully on the GPU. On smaller cards it may fail to load: lower num_ctx or remove num_gpu.
  • For the VRAM figure above, set these on the Ollama server:
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0

Usage

ollama run MobiusDevelopment/Qwen3.6-35B-A3B-A8E4-128K

OpenAI-compatible endpoint: http://localhost:11434/v1, model MobiusDevelopment/Qwen3.6-35B-A3B-A8E4-128K.

Default sampling: temperature 0.6, top_p 0.95, top_k 20, min_p 0.

Notes

  • Requantized from unsloth/Qwen3.6-35B-A3B-GGUF Q8_0 with llama-quantize --tensor-type, no imatrix.
  • KL divergence and perplexity were measured against Q8_0 on ~10K tokens of C++ and Python source.
  • No MTP heads, so speculative decoding does not apply.

License

Apache 2.0, inherited from Qwen3.6-35B-A3B.