1,885 Pulls 1 Tag Updated 1 month ago
217 Pulls 1 Tag Updated 1 month ago
Nanbeige4.1-3B illustrates that compact models can simultaneously achieve robust reasoning, preference alignment, and effective agentic behaviors
7,456 Pulls 6 Tags Updated 7 months ago
3B model that shouldn't be this good - crushes benchmarks through deep chain-of-thought reasoning
1,870 Pulls 1 Tag Updated 7 months ago
26.2.27. Nanbeige4.1-3B-Q4 illustrates that compact models can simultaneously achieve robust reasoning, preference alignment, and effective agentic behaviors.
868 Pulls 1 Tag Updated 6 months ago
Fine-tuned version of Nanbeige 4.1 3B specialized for Python code generation with direct, focused output.
823 Pulls 3 Tags Updated 7 months ago
Nanbeige4.1-3B-q4_K_M no think tools fit 4G-6G GPU OpenClaw local free tokens LobsterAI
475 Pulls 1 Tag Updated 6 months ago
Ollama version of de-censored Nanbeige4.1-3B-heretic
427 Pulls 1 Tag Updated 7 months ago
Optimized its system prompt and parameters for not to overthink too much.
343 Pulls 1 Tag Updated 7 months ago
nanbeige4.1-3b-tools no think fit 4G-6G GPU OpenClaw local free tokens
245 Pulls 1 Tag Updated 6 months ago
112 Pulls 1 Tag Updated 6 months ago
66 Pulls 1 Tag Updated 6 months ago
38 Pulls 1 Tag Updated 6 months ago
32 Pulls 1 Tag Updated 6 months ago
The Nanbeige2-16B-Chat is the latest 16B model developed by the Nanbeige Lab, which utilized 4.5T tokens of high-quality training data during the training phase.
150 Pulls 1 Tag Updated 2 years ago
Model şu an beta aşamasındadır. 1M parametrenin getirdiği sınırlardan dolayı karmaşık cümle yapılarında bozulmalar yaşanabilir. Geliştirme süreci devam etmektedir.
28 Pulls 3 Tags Updated 8 months ago
Instinct is Continue's state-of-the-art open Next Edit model. Robustly fine-tuned from Qwen2.5-Coder-7B, Instinct intelligently predicts your next move to keep you in flow.
6,939 Pulls 1 Tag Updated 1 year ago
A 62M-parameter GPT trained from scratch on a single 8GB-RAM NVIDIA Jetson, offering a compact open-weights base model for raw text completion. Not instruction-tuned.
4 Pulls 1 Tag Updated 1 month ago
Llama-3.1-Nemotron-70B-Instruct is a large language model customized by NVIDIA to improve the helpfulness of LLM generated responses to user queries.
3,150 Pulls 6 Tags Updated 1 year ago
TGAI NB — a lightweight Chinese MoE chat model family, built end-to-end by a high school student. V2 (3B): 8-expert sparse MoE, better quality for desktop. V1 (0.86B): 4-expert sparse MoE, ultra-light, smooth on low-end devices.
36 Pulls 3 Tags Updated 3 weeks ago