19 4 days ago

# POCKET-26B **On-device Korean AI, based on Google Gemma4-26B-A4B.** Runs on your PC or phone with **no GPU** — in Ollama, LM Studio, PocketPal, or any llama.cpp app. ## Run

tools thinking
ollama run VIDRAFT/pocket-26b

Applications

Claude Code
Claude Code ollama launch claude --model VIDRAFT/pocket-26b
OpenCode
OpenCode ollama launch opencode --model VIDRAFT/pocket-26b
Hermes Agent
Hermes Agent ollama launch hermes --model VIDRAFT/pocket-26b
OpenClaw
OpenClaw ollama launch openclaw --model VIDRAFT/pocket-26b

Models

View all →

Readme

What it is

POCKET-26B takes Google’s Gemma4-26B-A4B (25.2B total / ~4B active MoE, Apache-2.0) and re-quantizes it with VIDRAFT’s Korean-tuned mixed-precision quantization — kept unpruned, so quality holds. Because it’s Gemma4, it loads in every mainstream runtime today (Ollama, LM Studio, PocketPal, koboldcpp, browser) — no bleeding-edge build required.

Quality — GPQA-Diamond (greedy, 198 questions)

Build GPQA-Diamond vs base
Gemma4-26B-A4B (base) 67.7%
POCKET-26B (Q4_K_M) 67.7% = base (lossless)

Statistically lossless versus the base — at 17 GB, GPU-optional.

Details

  • Base model: google/gemma-4-26B-A4B-it (Apache-2.0)
  • Our work: Korean imatrix + mixed-precision quantization (no pruning)
  • Size: 17 GB (Q4_K_M) · Context: 262K
  • License: Apache-2.0

Also available

  • Full GGUF (Q4 / Q2) on Hugging Face: hf.co/FINAL-Bench/POCKET-26B-GGUF
  • Live CPU demo (no GPU): Hugging Face Spaces — FINAL-Bench/POCKET-26B-CPU

POCKET is a VIDRAFT on-device model family. Big models, small hardware — no GPU, no cloud.