DeepSeek-R1 is a family of open reasoning models with performance approaching that of leading models, such as O3 and Gemini 2.5 Pro.
92.2M Pulls 35 Tags Updated 1 year ago
Olmo is a series of Open language models designed to enable the science of language models. These models are pre-trained on the Dolma 3 dataset and post-trained on the Dolci datasets.
453.7K Pulls 15 Tags Updated 8 months ago
287.7K Pulls 10 Tags Updated 8 months ago
Qwen 3.8 Ollama profiles for RTX 5090 27B with vision, thinking mode, and native tool calling
406 Pulls 2 Tags Updated 2 weeks ago
Qwen 3.6 Ollama profiles for RTX 5090 across 27B dense and 35B-A3B MoE variants, with vision, thinking mode, and native tool calling.
813 Pulls 3 Tags Updated 2 months ago
An open 30B MoE model from NVIDIA with 3B activated parameters that delivers strong reasoning and agentic capabilities.
144.9K Pulls 3 Tags Updated 5 months ago
A new collection of open translation models built on Gemma 3, helping people communicate across 55 languages.
2.2M Pulls 13 Tags Updated 7 months ago
Atom-Olmo3-7B is a specialized language model fine-tuned from Olmo-3 7B Instruct for collaborative problem-solving and creative exploration.
222 Pulls 1 Tag Updated 9 months ago
Quantized version of Openhands-lm-32B-v01 optimized for tool usage in Cline / Roo Code and complex problem solving.
768 Pulls 3 Tags Updated 1 year ago
Ollama version of https://huggingface.co/defog/llama-3-sqlcoder-8b with full template and example system prompt usage as DDL Statement. Forked from https://ollama.com/mannix/defog-llama3-sqlcoder-8b:latest
364 Pulls 1 Tag Updated 1 year ago
Vidwan is a customized AI model based on LLaMA 3, designed to provide academic and research-based responses. It specializes in assisting with complex topics, answering in-depth questions, and providing structured insights.
81 Pulls 1 Tag Updated 1 year ago
A text-only, thinking-capable variant of Qwen3.5-35B-A3B — leaner and faster by removing the CLIP vision projector. Based on Unsloth's Q4_K_M quantization of Alibaba's Qwen3.5-35B-A3B.
2,272 Pulls 2 Tags Updated 5 months ago
This is the Q4_K converted version of the original Qiskit/mistral-small-3.2-24b-qiskit provided by IBM's Qiskit Huggingface. Please refer to the original mistral-small-3.2-24b-qiskit model card for more details.
30 Pulls 1 Tag Updated 2 months ago
A 7B math reasoning model from Allen AI, trained with RL-Zero to solve problems step-by-step like a skilled tutor. Supports 65K context for complex multi-step problems - runs on any laptop.
260 Pulls 7 Tags Updated 9 months ago
44 Pulls 1 Tag Updated 6 months ago
Adapted for Cline tool / Roo Code use in VS Code fused model , hybrid of DeepSeekR1 and Qwen2.5 coder, from FuseAI/FuseO1-DeepSeekR1-Qwen2.5-Coder-32B-Preview.
4,937 Pulls 2 Tags Updated 1 year ago
The quantized versions of the FuseO1 + DeepSeek R1 + QwQ + Sky T1 Fusion Model
520 Pulls 8 Tags Updated 1 year ago
Poro 34b chat is a chat-tuned version of Poro 34B trained to follow instructions in both Finnish and English.
268 Pulls 1 Tag Updated 1 year ago
MiniCPM-V surpasses proprietary models such as GPT-4V, Gemini Pro, Qwen-VL and Claude 3 in overall performance, and support multimodal conversation for over 30 languages.
45.9K Pulls 8 Tags Updated 2 years ago
34B parameter decoder-only transformer pretrained on Finnish, English and code.
942 Pulls 1 Tag Updated 2 years ago