Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
Qwen3 Max · Ollama
Search for models on Ollama.
  • mdq100/qwen3.5

    Custom Qwen3.5 variants optimized for 128GB unified memory systems, such as AMD Ryzen AI Max+ 395. On Windows 11, GPU is limited to 96GB (32GB reserved for OS/CPU), requiring context window capped at 131072 tokens (128K) to fit within GPU memory limits.

    vision tools thinking

    431  Pulls 2  Tags Updated  5 months ago

  • srchmnmichael/qwen3.5-9B-uncensored

    Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

    1,018  Pulls 4  Tags Updated  4 days ago

  • tinyrick/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-Vision

    vision

    2,255  Pulls 1  Tag Updated  4 weeks ago

  • n0404n0404/qwen3.6-finetune-qwen3.8-max-glm5.2-kimi-k3-distillation-a56-1168cb-heretic

    567  Pulls 4  Tags Updated  3 weeks ago

  • edtorre/qwen3.6-hermes

    Qwen 3.6 27B (Q4_K_M) optimized for Hermes Agent — 64K context, 8192 max tokens, MTP for speed, flash attention + Q8 KV cache.

    vision tools thinking

    810  Pulls 1  Tag Updated  1 month ago

  • Omoeba/qwen3-coder-max

    maximum 256k context length by default

    tools 30b

    1,535  Pulls 1  Tag Updated  11 months ago

  • Omoeba/qwen3-2507-abliterated-max

    maximum 256k context length for coding and other long-context tasks

    tools 30b

    871  Pulls 1  Tag Updated  11 months ago

  • aware/qwen3.6-40b-deck-opus-neo-code

    DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF

    1,733  Pulls 1  Tag Updated  3 months ago

  • zdolny/qwen3-coder58k-tools

    qwen3-coder with tools calling, context 58k to match full memory 31GB on RTX5090

    tools

    586  Pulls 1  Tag Updated  1 year ago

  • joachimhansen/opencode-qwen3

    Made for a rig with TRX 3080 10Gb, 32gb RAM and Ryzen 5800x

    tools

    31  Pulls 1  Tag Updated  2 months ago

  • S00K/alice

    Alice is a roleplay LLM based on qwen3:4b-instruct-2507-q4_K_M, simulating Alice Marie Chen, a 20yo from Portland. Features an 8k context window and is optimized for WhatsApp. The persona is an anti-assistant with a texting style, refusing technical tasks

    tools thinking

    705  Pulls 1  Tag Updated  7 months ago

  • lucloner/wen36-35b-uncensored-1m

    Qwen3.6-35B-A3B Uncensored is a Mixture-of-Experts model with 35.5B total parameters and 3B activated, extended to 1M-token context with the official MTP speculative-decoding layer, delivering 1.5x decode speedup on Ollama, with native thinking and tools.

    tools thinking

    347  Pulls 2  Tags Updated  2 weeks ago

  • rafw007/qwen36-a3b-claude-coder

    Qwen3.6-35B-A3B MoE coding agent for Claude Code / Codex / opencode, 64K context, native tool-calling, honest tool use, safety guardrails intact

    vision tools thinking

    3,559  Pulls 2  Tags Updated  2 months ago

  • Havenlon/Execution-Boundary-Qwen35-27B-Q4_K_M

    A Havenlon-focused Qwen3.5 27B model for deep reasoning about execution boundaries, Adversarial Completeness, AI Agent control, evidence chains, and real-world execution.

    2  Pulls 3  Tags Updated  1 week ago

  • ExpedientFalcon/qwen3-14b-agent-m2max

    tools thinking

    96  Pulls 1  Tag Updated  1 year ago

  • srchmnmichael/Qwen3.8-Uncensored

    Qwen3.8-27B tensor-level abliterated, vision tower and MTP head untouched. 0% over-refusal on XSTest, 0-6% refusal across the A/B suite, no measurable capability loss. Full mmproj vision, tool calling and thinking, 262K context.

    vision tools thinking

    2,064  Pulls 4  Tags Updated  4 days ago

  • orcarouter/Qwen3.8-27B-Uncensored

    Qwen3.8-27B tensor-level abliterated, vision tower and MTP head untouched. 0% over-refusal on XSTest, 0-6% refusal across the A/B suite, no measurable capability loss. Full mmproj vision, tool calling and thinking, 262K context.

    vision tools thinking

    129.6K  Pulls 20  Tags Updated  4 days ago

  • oamazonasgabriel/qwen3.6-35b-a3b

    A memory-efficient model configuration of Qwen3.6-35B-A3B using an upstream imatrix-calibrated IQ4_XS quantization and q4_0 KV cache. Designed for 24 GB VRAM

    tools thinking

    2,022  Pulls 1  Tag Updated  2 months ago

  • fredrezones55/Qwen3.5-APEX

    Qwen3.5-35B-A3B APEX GGUF -- A Novel MoE-Aware Mixed-Precision Quantization Technique Brought to you by the LocalAI team -- the creators of LocalAI the open-source AI engine that runs any model - LLMs, vision, image - on any hardware.

    vision tools thinking

    755  Pulls 2  Tags Updated  4 months ago

  • Maternion/manim-coder

    Fine-tuned manim model using Qwen2.5-Coder-14B-Instruct trained on 3blue1brown-manim dataset with 2,407 examples.

    tools 14b

    541  Pulls 3  Tags Updated  6 months ago

© 2026 Ollama
Blog Support