Ollama
Models Docs Pricing
Sign in Download
Models Download Docs Pricing Sign in
⇅
GLM-4.7 · Ollama
Search for models on Ollama.
  • glm-4.7-flash

    As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

    tools thinking

    1.5M  Pulls 4  Tags Updated  2 months ago

  • dlasher/GLM-4.7-Flash-GGUF

    12  Pulls 1  Tag Updated  2 days ago

  • byczech/glm-4.7-flash-uncensored-16G-AU-IQ3_M

    Uncensored GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_M and optimized for 16 GB VRAM. Suitable for coding, agents, automation, roleplay, analysis and unrestricted local experimentation.

    tools thinking 30b

    1,308  Pulls 1  Tag Updated  1 month ago

  • rafw007/glm-4.7-flash-opencode

    A family of custom models built on **GLM-4.7-Flash** (MoE, 30B total / 3B active), tuned to act as autonomous coding agents — **each variant targeting a specific harness**

    tools thinking

    604  Pulls 1  Tag Updated  2 months ago

  • byczech/glm-4.7-flash-16G-UD-Q3_K_XL

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to Q3_K_XL and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 136K fitting within 16 GB in suitable configurations.

    tools thinking 30b

    202  Pulls 1  Tag Updated  1 month ago

  • huihui_ai/glm-4.7-flash-abliterated

    As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

    tools thinking

    292.4K  Pulls 5  Tags Updated  7 months ago

  • byczech/glm-4.7-flash-12G-UD-IQ2_XSS

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.

    tools thinking 30b

    186  Pulls 1  Tag Updated  1 month ago

  • byczech/glm-4.7-flash-16G-UD-IQ3_XSS

    GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_XSS and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 180K fitting within 16 GB in suitable configurations.

    tools thinking 30b

    94  Pulls 1  Tag Updated  1 month ago

  • moophlo/GLM-4.7-Flash-GGUF

    113  Pulls 2  Tags Updated  5 months ago

  • tk2133/GLM-4.7-Flash-REAP-23B-A3B

    Unsloth/GLM-4.7-Flash-REAP-23B-A3B model

    tools thinking

    1,985  Pulls 17  Tags Updated  7 months ago

  • SimonPu/GLM-4.7-Flash

    This model was base on unsloth/GLM-4.7-Flash and trained on a small reasoning dataset of Claude Opus 4.5, with reasoning effort set to High.

    tools thinking

    1,190  Pulls 4  Tags Updated  6 months ago

  • aia/GLM-4.7-Flash-REAP-23B-A3B-GGUF

    A memory-efficient compressed variant of GLM-4.7-Flash that maintains near-identical performance while being 25% lighter.

    855  Pulls 1  Tag Updated  7 months ago

  • DedeProgames/codex-oss

    An open-source model based on GLM-4.7-Flash, optimized specifically for code generation and development workflows.

    tools thinking

    622  Pulls 1  Tag Updated  6 months ago

  • MrScratchcat22/GLM-4.7-Flash-REAP-23B-A3B

    tools thinking

    501  Pulls 1  Tag Updated  7 months ago

  • aratan/GLM-4.7-Agent-Flash

    tools thinking

    105  Pulls 1  Tag Updated  6 months ago

  • siddhjain68/GLM_4.7

    tools thinking cloud

    85  Pulls 1  Tag Updated  6 months ago

  • alexmoneo777/glm-4.7-flash

    Glm-4.7-flash

    tools

    56  Pulls 1  Tag Updated  5 months ago

  • bergencvv/cerebras-glm-4.7-reap-218b-a32b

    13  Pulls 1  Tag Updated  4 weeks ago

  • dhiltgen/glm-4.7-flash

    tools thinking

    193  Pulls 10  Tags Updated  1 month ago

  • kalikorpz/glm-4.7-flash

    tools

    181  Pulls 1  Tag Updated  4 months ago

© 2026 Ollama
Blog Support