3,384 2 months ago

tools thinking
ollama run batiai/gemma4-12b:q3

Details

2 months ago

d04fc1e15a2b · 6.1GB ·

gemma4
·
11.9B
·
Q3_K_M
You are a helpful AI assistant.
{ "num_ctx": 131072, "stop": [ "<turn|>", "<eos>" ], "temperature":

Readme

Gemma 4 12B-it — Quantized by BatiAI

Google DeepMind’s encoder-free multimodal model — text + image + audio + video, running on a 16GB Mac. 26B-MoE-class quality at official Google weights.

Models

Tag Size RAM target Use Case
q2 ~4.2GB 8GB Mac Ultra-compact
iq3 ~4.6GB 8GB Mac imatrix, smallest
q3 ~5.7GB 8GB+ Mac Balanced
iq4 ~6.2GB 16GB Mac imatrix, best size/quality
q4 ~6.9GB 16GB Mac Recommended
q6 ~9.2GB 16GB+ Near-original

Quick Start

ollama run batiai/gemma4-12b:q4

Why Gemma 4 12B?

  • 26B-MoE-class quality at
  • Encoder-free multimodal — first mid-sized model with native audio. No separate vision/audio encoder
  • Vision+reasoning: DocVQA 94.9 · InfoVQA 88.4 · MMMU-Pro 69.1 · AIME 2026 77.5
  • 256K context, 140+ languages
  • Apache 2.0
  • Released June 3, 2026

RAM Requirements

Your Mac RAM q2 iq3 q3 iq4 q4 q6
8GB ⚠️
16GB
24GB+

Multimodal (image / audio / video)

The Ollama tags above are the text LLM. For vision/audio/video, use the mmproj file with llama.cpp:

llama-mtmd-cli -m gemma-4-12B-it-Q4_K_M.gguf \
  --mmproj mmproj-google-gemma-4-12B-it-BF16.gguf --image photo.jpg -p "Describe this."

(mmproj is on the HF repo — encoder-free so it’s only ~170MB.)

For Other Macs

Your Mac Recommended
8GB batiai/gemma4-12b:q3 or batiai/gemma4-e2b:q4
16GB batiai/gemma4-12b:q4 (this, recommended)
24GB batiai/gemma4-26b:iq4
32GB batiai/nemotron3-nano:iq4

Why BatiAI?

  • Quantized directly from official Google weights
  • imatrix calibrated (IQ variants)
  • Low quants (q2/q3) for 8GB Macs
  • Vision+audio mmproj included
  • BatiAI signed (general.author=BatiAI)

License

Apache 2.0 — quantized from google/gemma-4-12B-it.

Built for BatiFlow

Free, on-device AI automation for Mac. 5MB app, 100% local, unlimited.

https://flow.bati.ai