20 4 days ago

GLM-4.7 Flash 30B text model with tool-calling support, quantized to Q3_K_XL and optimized for 16 GB VRAM. The model supports up to 202,752 tokens context window, with approximately 136K fitting within 16 GB in suitable configurations.

tools thinking 30b
543c6a8262f9 · 63B
{
"min_p": 0.01,
"repeat_penalty": 1,
"temperature": 1,
"top_p": 0.95
}