Gemma 4 31B multimodal instruct model with vision support, quantized to Q3_K_S and optimized for 16 GB VRAM. Supports a practical context window of approximately 36K–40K tokens, depending on the Ollama version, GPU, backend and runtime configuration.