-
glm-4.7-flash-uncensored-16G-AU-IQ3_M
Uncensored GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_M and optimized for 16 GB VRAM. Suitable for coding, agents, automation, roleplay, analysis and unrestricted local experimentation.
tools thinking 30b237 Pulls 1 Tag Updated 3 days ago
-
qwen3.6-mmproj-16G-UD-IQ3_S
Qwen 3.6 35B multimodal model with vision support, quantized to IQ3_S and optimized to fit within 16 GB VRAM with Q4_0 KV cache. Supports very large context windows, with the practical maximum depending on the Ollama version, GPU and backend.
vision tools thinking 35b168 Pulls 1 Tag Updated 1 week ago
-
qwen3.6-16G-UD-IQ3_S
Qwen 3.6 35B text-only model, quantized to IQ3_S and optimized to fit within 16 GB VRAM with Q4_0 KV cache. Designed for long-context chat, coding, document analysis and general-purpose local inference. Supports very large context windows.
tools thinking 35b98 Pulls 1 Tag Updated 1 week ago
-
gemma4-it-mmproj-16G-Q3_K_S
Gemma 4 31B multimodal instruct model with vision support, quantized to Q3_K_S and optimized for 16 GB VRAM. Supports a practical context window of approximately 36K–40K tokens, depending on the Ollama version, GPU, backend and runtime configuration.
vision tools thinking 31b84 Pulls 1 Tag Updated 1 week ago
-
glm-4.7-flash-12G-UD-IQ2_XSS
GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.
tools thinking 30b29 Pulls 1 Tag Updated 3 days ago
-
glm-4.7-flash-16G-UD-Q3_K_XL
GLM-4.7 Flash 30B text model with tool-calling support, quantized to Q3_K_XL and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 136K fitting within 16 GB in suitable configurations.
tools thinking 30b20 Pulls 1 Tag Updated 3 days ago
-
glm-4.7-flash-16G-UD-IQ3_XSS
GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ3_XSS and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 180K fitting within 16 GB in suitable configurations.
tools thinking 30b5 Pulls 1 Tag Updated 3 days ago