Agentic coding model for 24 GB GPUs: TeichAI's Gemma-4-31B Fable-5 agent distill (vision + thinking + tools) with a disciplined coding-agent system prompt baked in. Inspect → reproduce → smallest fix → re-verify.
149 Pulls 1 Tag Updated 1 month ago
Generic all purpose model. Occasionally may have notable logic, usually Llama-3_3-Nemotron-Super-49B-v1_5 is preferred.
133 Pulls 1 Tag Updated 9 months ago
129 Pulls 1 Tag Updated 9 months ago
Gemma 4 E4B (Google DeepMind) with thinking mode disabled. Compact multimodal model — 4.5B effective / 8B total parameters. Supports text, image and audio input. Designed for edge devices and local deployment. Knowledge cutoff: January 2025.
1,645 Pulls 1 Tag Updated 4 months ago
Gemma 4 E4B (Google DeepMind) with thinking mode enabled. Compact multimodal model — 4.5B effective / 8B total parameters. Supports text, image and audio input. Designed for edge devices and local deployment. Knowledge cutoff: January 2025.
844 Pulls 1 Tag Updated 4 months ago
Gemma 4 Ollama profiles for RTX 4090/5090 across 12B, 26B-A4B, and 31B variants, with multimodal support and native tool calling
477 Pulls 5 Tags Updated 2 months ago
Parable is IBM Granite 4.1 fine-tuned on Claude Fable 5 and GPT-5.5 agent traces. Adds think reasoning to Granite. Sibling: Qwen3 line at parable/qwen3-fable
393 Pulls 10 Tags Updated 1 month ago
Gemma 4 distilled from claude opus 4.6 thinking. Has only a 5% gap with claude opus 4.6 thinking while being over 40x smaller. Designed for server inference. Designed for local inference
629 Pulls 1 Tag Updated 4 months ago
A PORT TO OLLAMA FROM THE ORIGINAL: https://huggingface.co/KatyTheCutie/LemonadeRP-4.5.3
288 Pulls 2 Tags Updated 1 year ago