-
Ternary-Bonsai-2-27B
Mirror of Prism ML's Bonsai 2 27B ternary GGUF packs (PQ2_0 and PTQ1_0): a 27B hybrid-attention reasoning model in 6.70 GiB or 5.53 GiB, 262K context, vision and tool calling, Apache-2.0. Requires Prism ML's llama.cpp fork — stock Ollama cannot load
vision tools thinking 27b1,674 Pulls 11 Tags Updated yesterday
-
qwen3.8-9b-distill
Qwen3.8-9B-Distill is a dense 9B reasoning model: it brings the reasoning behaviour of a frontier-scale teacher (Qwen3.8, 2.4T-A95B) into a model that fits on a single consumer GPU. Every answer opens with a <think> block learned from real teacher traces,
vision tools thinking1,015 Pulls 3 Tags Updated 6 days ago
-
minicpm5-2b
Compact 2.5B model by OpenBMB — 2B-class open-source SOTA, built for on-device, local assistants, coding agents and tool-use workflows
283 Pulls 5 Tags Updated 1 week ago
-
oxcoder-9b
OxCoder-9B by OrionLLM: 9B coding model distilled from frontier agent traces (Claude Code / OpenCode / Codex). 262K context, native tool calling, vision. Q4_K_M, Q5_K_M.
vision tools thinking179 Pulls 2 Tags Updated 6 days ago
-
nex-n2.5-mini
Nex-N2.5-mini (nex-agi) : MoE 35B / ~3B actifs post-trained on Qwen3.5-35B-A3B. Vision, tool-calling, reasoning, 256k context. Quants Q5_K_M, Q4_K_M (latest), Q3_K_M, Q2_K.
vision tools thinking168 Pulls 4 Tags Updated 6 days ago