348 Downloads Updated 1 month ago
ollama run pdurlej/gemma-4-26B-A4B-it-heretic
ollama launch claude --model pdurlej/gemma-4-26B-A4B-it-heretic
ollama launch opencode --model pdurlej/gemma-4-26B-A4B-it-heretic
ollama launch hermes --model pdurlej/gemma-4-26B-A4B-it-heretic
ollama launch openclaw --model pdurlej/gemma-4-26B-A4B-it-heretic
A fast, capable and reduced-refusal Gemma 4 MoE for creative work, coding and unrestricted local exploration.
This reduced-refusal variant of Gemma 4 26B-A4B uses a Mixture-of-Experts architecture, activating only part of the model for each token to balance capability and generation speed.
| Architecture | Gemma 4 Mixture-of-Experts |
| Quantization | Q4_K_M |
| Download size | 17 GB |
| Context | Up to 256K tokens |
| Input | Text |
| Capabilities | Thinking and tools |
ollama run pdurlej/gemma-4-26B-A4B-it-heretic:Q4_K_M
Start with an 8K–16K context window and increase it gradually. The full 256K context requires considerably more memory than the model weights alone.
| Usage | Practical recommendation |
|---|---|
| Minimum | Apple Silicon MacBook with 24GB unified memory; close heavy apps and use 4K–8K context |
| Recommended | MacBook Pro with 32GB–48GB; comfortable use at approximately 16K–32K context |
| Larger contexts and multitasking | MacBook Pro with 64GB or more unified memory |
Allow approximately 22GB of free storage. A MacBook Air can run the model, but a MacBook Pro offers active cooling and higher sustained performance.
This model has reduced refusal behavior. Apply your own safety, privacy and content controls when using it in applications or sharing its output.
Original model: coder3101/gemma-4-26B-A4B-it-heretic