Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
Qwen3.8 Max · Ollama
Search for models on Ollama.
  • edtorre/qwen3.6-hermes

    Qwen 3.6 27B (Q4_K_M) optimized for Hermes Agent — 64K context, 8192 max tokens, MTP for speed, flash attention + Q8 KV cache.

    vision tools thinking

    373  Pulls 1  Tag Updated  3 weeks ago

  • mdq100/qwen3.5

    Custom Qwen3.5 variants optimized for 128GB unified memory systems, such as AMD Ryzen AI Max+ 395. On Windows 11, GPU is limited to 96GB (32GB reserved for OS/CPU), requiring context window capped at 131072 tokens (128K) to fit within GPU memory limits.

    vision tools thinking

    384  Pulls 2  Tags Updated  4 months ago

  • zdolny/qwen3-coder58k-tools

    qwen3-coder with tools calling, context 58k to match full memory 31GB on RTX5090

    tools

    572  Pulls 1  Tag Updated  11 months ago

  • rafw007/qwen3-coder-next-80b-redteam

    Validated runtime configuration + measured results for an abliterated (uncensored) Qwen3-Coder-Next 80B-A3B running as an agentic coding backend on CPU via ik_llama.cpp

    279  Pulls 1  Tag Updated  11 hours ago

© 2026 Ollama
Blog Contact