84 1 week ago

Gemma 4 31B multimodal instruct model with vision support, quantized to Q3_K_S and optimized for 16 GB VRAM. Supports a practical context window of approximately 36K–40K tokens, depending on the Ollama version, GPU, backend and runtime configuration.

vision tools thinking 31b
56380ca2ab89 · 42B
{
"temperature": 1,
"top_k": 64,
"top_p": 0.95
}