202 Pulls 1 Tag Updated 3 days ago
# POCKET-26B **On-device Korean AI, based on Google Gemma4-26B-A4B.** Runs on your PC or phone with **no GPU** — in Ollama, LM Studio, PocketPal, or any llama.cpp app. ## Run
33 Pulls 1 Tag Updated 1 week ago
Gemma 4 26B A4B Instruct GGUF Q4_K_M build for Ollama, configured with a 256K context window. This is a text-focused local model built from google/gemma-4-26B-A4B-it and intended for chat, summarization, tool-style workflows, and long-context testing.
116.9K Pulls 1 Tag Updated 1 month ago
100 Pulls 1 Tag Updated 2 weeks ago
Gemma 4 26B (IQ4_XS) - Optimized for 16GB VRAM
21.4K Pulls 1 Tag Updated 4 months ago
Gemma 4 Uncensored 26B (IQ4_XS) - Optimized for 16GB VRAM
8,511 Pulls 1 Tag Updated 3 months ago
Gemma 4 26B MoE quantized by BatiAI. 77 t/s on M4 Max. Requires 24GB+ Mac.
6,678 Pulls 6 Tags Updated 4 months ago
Gemma 4 26B-A4B uncensored MoE: 1M context, vision, fast. ~91% recall, honestly documented.
1,873 Pulls 1 Tag Updated 1 month ago
Gemma 4 26B MoE (Google DeepMind) with thinking mode enabled. Mixture-of-Experts — 25.2B total / 3.8B active parameters, 256K context. Supports text and image input. Knowledge cutoff: January 2025.
2,756 Pulls 1 Tag Updated 4 months ago
Gemma 4 26B Optimized for 16GB VRAM via Q3 Quantization
1,075 Pulls 2 Tags Updated 3 months ago
Niestandardowy model Gemma 4 26B (~25,8B parametrów), dostrojony do działania jako niezależny agent kodowania i administracji . Obsługuje API zgodne z Anthropic, dzięki czemu obsługuje Claude Code, Codex i Opencode
740 Pulls 1 Tag Updated 2 months ago
Gemma 4 26B MoE (Google DeepMind) with thinking mode disabled. Mixture-of-Experts — 25.2B total / 3.8B active parameters, 256K context. Supports text and image input. Knowledge cutoff: January 2025.
1,035 Pulls 1 Tag Updated 4 months ago
Gemma 4 Ollama profiles for RTX 4090/5090 across 12B, 26B-A4B, and 31B variants, with multimodal support and native tool calling
428 Pulls 5 Tags Updated 2 months ago
Gemma 4 26B tuned for agentic use. 64K context window, flash attention + Q8 KV cache quantization for reduced VRAM. Temperature 0.7, output capped at 8192 tokens.
193 Pulls 1 Tag Updated 1 month ago
😈 Uncensored Gemma 4 26B (A4B MoE) for Ollama. 👁️ Native Vision, 🛠️ Tool Calling, GGUF, Rust, Linux, Windows AI Coding Agents & Multimodal Local AI Development. 🚀
1,032 Pulls 1 Tag Updated 3 weeks ago
Gemma 4 abliterated Quants (from https://huggingface.co/jenerallee78/gemma-4-26B-A4B-it-ara-abliterated)
8,138 Pulls 6 Tags Updated 4 months ago
Gemma-4-26B-A4B UD-Q8_K_XL MTP from Unsloth
164 Pulls 1 Tag Updated 1 month ago
llmfan46/gemma-4-26B-A4B-it-uncensored-heretic - quantized to q4_K_M from HF with vision capability retained
5,009 Pulls 1 Tag Updated 3 months ago
Huihui4-8B-A4B is a lightweight MoE (Mixture of Experts) conversational model optimized from Google's gemma-4-26B-A4B-it architecture
2,863 Pulls 4 Tags Updated 3 months ago
mradermacher/gemma-4-26B-A4B-it-heretic-GGUF
2,289 Pulls 1 Tag Updated 4 months ago