Data & AI Engineering - Open Source, Private AI Enthusiast @ impacte.tech
-
qwen3.6-35b-a3b
A memory-efficient model configuration of Qwen3.6-35B-A3B using an upstream imatrix-calibrated IQ4_XS quantization and q4_0 KV cache. Designed for 24 GB VRAM
tools thinking2,091 Pulls 1 Tag Updated 3 months ago
-
lfm2.5-2.6b
LFM2.5-2.6B by Liquid AI — deploy agents everywhere. A 2.6B dense reasoning model (Q4_K_M, ~1.7 GB) with native tool calling, 128K context, and 16-language support. Runs on any 8 GB GPU.
1,564 Pulls 1 Tag Updated 1 month ago
-
qwen3.5-9b
A coding-optimized configuration of Qwen3.5-9B designed for 16 GB single-GPU hardware. The model uses the official Q4_K_M quantization (~6.6 GB weights), leaving ~9 GB headroom for KV cache — enabling 32K+ context windows comfortably.
vision tools thinking1,043 Pulls 1 Tag Updated 3 months ago
-
qwen3.8-27b
Qwen3.8-27B in Q4_K_M quantization (32GB+ VRAM required). Dense 27.8B parameters with hybrid attention for long context (256K tokens). Apache 2.0 license. Ideal for high-VRAM setups (RTX 5090/4090 dual, M4 Ultra, etc.).
939 Pulls 4 Tags Updated 1 week ago
-
lfm2.5-230m
LFM2.5-230M is a hybrid language model by Liquid AI, Built to Run Anywhere. The ideal lightweight AI companion for: 4 GB laptops · 8 GB desktops · Edge devices · Raspberry Pi 5
399 Pulls 1 Tag Updated 5 days ago
-
qwen3.5-4b
A balanced configuration variant of Qwen3.5-4B, planned for FIM (Fill-In-the-Middle) and Tool Calling in restrict capacity environments. Using the Q3_K_S GGUF quantization from HuggingFace (unsloth).
398 Pulls 2 Tags Updated 1 month ago
-
qwen3.5-0.8b
Qwen3.5 0.8B in Q8_0 quantization. Small, fast model with 262K context, vision, tool use, and thinking capabilities. Optimized for local/edge deployment on constrained hardware.
vision tools thinking361 Pulls 1 Tag Updated 1 month ago
-
bonsai-27b
227 Pulls 1 Tag Updated 5 days ago
-
qwen3.5-2b
Qwen3.5 2B in Q8_0 quantization. Strong balance of capability and efficiency with 262K context, vision, tool use, and thinking. Ideal for local deployment on consumer hardware.
vision tools thinking219 Pulls 1 Tag Updated 1 month ago
-
qwen2.5-coder.1.5b-mlx
Code-specialized 1.5B model in F16 precision, optimized for MLX workflows. 32K context, fill-in-the-middle support, and fast inference on GPUs with 8 GB+ memory. Ideal for code generation, completion, and bug fixing.
199 Pulls 1 Tag Updated 1 month ago
-
qwen2.5-coder-0.5b
A lightweight, FIM (Fill-In-the-Middle) optimized variant of Qwen2.5-Coder-0.5B-Instruct using the f16 GGUF quantization from HuggingFace. At only ~1 GB, it fits comfortably on any 8 GB single GPU with headroom for 8K context.
162 Pulls 1 Tag Updated 1 month ago
-
nemotron-nano-9b-v2
Q4_K_M and BF16 quantizations of NVIDIA Nemotron Nano 9B v2, NVIDIA’s open 9B reasoning model. Q4 quantized locally from the BF16 source and tuned for a single 16 GB GPU card with maximum KV cache.
tools thinking128 Pulls 3 Tags Updated 3 weeks ago
-
lfm2.5-8b-a1b
A general-purpose 8.3B Mixture-of-Experts model from Liquid AI that activates only ~1.5B parameters per token — delivering strong reasoning, tool calling, and multilingual support while fitting comfortably in 8 GB VRAM.
112 Pulls 1 Tag Updated 1 month ago
-
nemotron-3.5-lightning
An open 30B MoE model with ~3B active parameters, packaged by impacte.tech for the execution layer of always-on agents. Uses the official Q4_K_M / IQ4_XS quantizations. Features native tool calling and thinking modes.
tools thinking95 Pulls 2 Tags Updated 2 weeks ago
-
lfm2-1.2b-tool
LFM2-1.2B-Tool — a specialized 1.2B tool-calling model from Liquid AI, fine-tuned exclusively for concise and precise function calling. Uses a hybrid LIV+GQA architecture, outperforms thinking models of similar size without chain-of-thought overhead.
67 Pulls 1 Tag Updated 2 months ago
-
ternary-bonsai-8btools thinking
33 Pulls 2 Tags Updated 5 days ago