27 3 days ago

GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.

tools thinking 30b
b507b9c2f6ca · 13B
{{ .Prompt }}