27 3 days ago

GLM-4.7 Flash 30B text model with tool-calling support, quantized to IQ2_XSS and optimized for 12 GB VRAM. Supports a maximum context window of 202,752 tokens, subject to the Ollama version, GPU, backend and runtime configuration.

tools thinking 30b
543c6a8262f9 · 63B
{
"min_p": 0.01,
"repeat_penalty": 1,
"temperature": 1,
"top_p": 0.95
}