Custom Qwen3.5 variants optimized for 128GB unified memory systems, such as AMD Ryzen AI Max+ 395. On Windows 11, GPU is limited to 96GB (32GB reserved for OS/CPU), requiring context window capped at 131072 tokens (128K) to fit within GPU memory limits.
431 Pulls 2 Tags Updated 5 months ago
Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
1,021 Pulls 4 Tags Updated 4 days ago
maximum 256k context length for coding and other long-context tasks
871 Pulls 1 Tag Updated 11 months ago
A Havenlon-focused Qwen3.5 27B model for deep reasoning about execution boundaries, Adversarial Completeness, AI Agent control, evidence chains, and real-world execution.
2 Pulls 3 Tags Updated 1 week ago
qwen3-coder with tools calling, context 58k to match full memory 31GB on RTX5090
586 Pulls 1 Tag Updated 1 year ago
Qwen3.5-35B-A3B APEX GGUF -- A Novel MoE-Aware Mixed-Precision Quantization Technique Brought to you by the LocalAI team -- the creators of LocalAI the open-source AI engine that runs any model - LLMs, vision, image - on any hardware.
755 Pulls 2 Tags Updated 4 months ago
Qwen3.6-35B-A3B Uncensored is a Mixture-of-Experts model with 35.5B total parameters and 3B activated, extended to 1M-token context with the official MTP speculative-decoding layer, delivering 1.5x decode speedup on Ollama, with native thinking and tools.
347 Pulls 2 Tags Updated 2 weeks ago
Qwen3.6-35B-A3B MoE coding agent for Claude Code / Codex / opencode, 64K context, native tool-calling, honest tool use, safety guardrails intact
3,559 Pulls 2 Tags Updated 2 months ago
A memory-efficient model configuration of Qwen3.6-35B-A3B using an upstream imatrix-calibrated IQ4_XS quantization and q4_0 KV cache. Designed for 24 GB VRAM
2,022 Pulls 1 Tag Updated 2 months ago