byczech/glm-4.7-flash-12G-UD-IQ2_XSS
GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.