126 2 months ago

Gemma 4 31B text-only instruct model with tool-calling and thinking support, quantized to Q3_K_S and optimized for 16 GB VRAM. Supports up to 256K context, with approximately 64K tokens fitting fully within 16 GB VRAM in tested configurations.

tools thinking 31b
56380ca2ab89 ยท 42B
{
"temperature": 1,
"top_k": 64,
"top_p": 0.95
}