Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
GLM-4.7 Flash · Ollama
Search for models on Ollama.
  • glm-4.7-flash

    As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

    tools thinking

    1.5M  Pulls 4  Tags Updated  2 months ago

  • dlasher/GLM-4.7-Flash-GGUF

    16  Pulls 1  Tag Updated  4 days ago

  • byczech/glm-4.7-flash-uncensored-16G-AU-IQ3_M

    Uncensored GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_M and optimized for 16 GB VRAM. Suitable for coding, agents, automation, roleplay, analysis and unrestricted local experimentation.

    tools thinking 30b

    1,340  Pulls 1  Tag Updated  1 month ago

  • rafw007/glm-4.7-flash-opencode

    A family of custom models built on **GLM-4.7-Flash** (MoE, 30B total / 3B active), tuned to act as autonomous coding agents — **each variant targeting a specific harness**

    tools thinking

    614  Pulls 1  Tag Updated  3 months ago

  • byczech/glm-4.7-flash-16G-UD-Q3_K_XL

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to Q3_K_XL and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 136K fitting within 16 GB in suitable configurations.

    tools thinking 30b

    217  Pulls 1  Tag Updated  1 month ago

  • huihui_ai/glm-4.7-flash-abliterated

    As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

    tools thinking

    297.5K  Pulls 5  Tags Updated  7 months ago

  • byczech/glm-4.7-flash-12G-UD-IQ2_XSS

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.

    tools thinking 30b

    191  Pulls 1  Tag Updated  1 month ago

  • byczech/glm-4.7-flash-16G-UD-IQ3_XSS

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_XSS and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 180K fitting within 16 GB in suitable configurations.

    tools thinking 30b

    101  Pulls 1  Tag Updated  1 month ago

  • moophlo/GLM-4.7-Flash-GGUF

    114  Pulls 2  Tags Updated  5 months ago

  • tk2133/GLM-4.7-Flash-REAP-23B-A3B

    Unsloth/GLM-4.7-Flash-REAP-23B-A3B model

    tools thinking

    1,990  Pulls 17  Tags Updated  7 months ago

  • SimonPu/GLM-4.7-Flash

    This model was base on unsloth/GLM-4.7-Flash and trained on a small reasoning dataset of Claude Opus 4.5, with reasoning effort set to High.

    tools thinking

    1,190  Pulls 4  Tags Updated  6 months ago

  • aia/GLM-4.7-Flash-REAP-23B-A3B-GGUF

    A memory-efficient compressed variant of GLM-4.7-Flash that maintains near-identical performance while being 25% lighter.

    856  Pulls 1  Tag Updated  7 months ago

  • DedeProgames/codex-oss

    An open-source model based on GLM-4.7-Flash, optimized specifically for code generation and development workflows.

    tools thinking

    628  Pulls 1  Tag Updated  6 months ago

  • MrScratchcat22/GLM-4.7-Flash-REAP-23B-A3B

    tools thinking

    502  Pulls 1  Tag Updated  7 months ago

  • aratan/GLM-4.7-Agent-Flash

    tools thinking

    105  Pulls 1  Tag Updated  6 months ago

  • alexmoneo777/glm-4.7-flash

    Glm-4.7-flash

    tools

    57  Pulls 1  Tag Updated  5 months ago

  • dhiltgen/glm-4.7-flash

    tools thinking

    194  Pulls 10  Tags Updated  1 month ago

  • kalikorpz/glm-4.7-flash

    tools

    182  Pulls 1  Tag Updated  4 months ago

  • gag0/glm-4.7-flash

    tools thinking

    33  Pulls 1  Tag Updated  5 months ago

  • sparksammy/glm-4.7-flash-unsloth

    tools thinking

    1,219  Pulls 5  Tags Updated  6 months ago

© 2026 Ollama
Blog Support