Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
GLM-4 FlashX · Ollama
Search for models on Ollama.
  • byczech/glm-4.7-flash-16G-UD-Q3_K_XL

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to Q3_K_XL and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 136K fitting within 16 GB in suitable configurations.

    tools thinking 30b

    218  Pulls 1  Tag Updated  1 month ago

  • byczech/glm-4.7-flash-12G-UD-IQ2_XSS

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.

    tools thinking 30b

    192  Pulls 1  Tag Updated  1 month ago

  • byczech/glm-4.7-flash-16G-UD-IQ3_XSS

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_XSS and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 180K fitting within 16 GB in suitable configurations.

    tools thinking 30b

    101  Pulls 1  Tag Updated  1 month ago

© 2026 Ollama
Blog Support