byczech/glm-4.7-flash-16G-UD-Q3_K_XL
GLM-4.7 Flash 30B text model with tool-calling support, quantized to Q3_K_XL and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 136K fitting within 16 GB in suitable configurations.