SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters.
4M Pulls 49 Tags Updated 1 year ago
🪐 A family of small models with 135M, 360M, and 1.7B parameters, trained on a new high-quality dataset.
2.1M Pulls 94 Tags Updated 2 years ago
Olmo is a series of Open language models designed to enable the science of language models. These models are pre-trained on the Dolma 3 dataset and post-trained on the Dolci datasets.
453.9K Pulls 15 Tags Updated 8 months ago
287.9K Pulls 10 Tags Updated 8 months ago
New state of the art 70B model. Llama 3.3 70B offers similar performance compared to the Llama 3.1 405B model.
4.1M Pulls 14 Tags Updated 1 year ago
A strong multi-lingual general language model with competitive performance to Llama 3.
1.2M Pulls 32 Tags Updated 2 years ago
Thinking model - efficient
44 Pulls 1 Tag Updated 2 months ago
SmolLM3 is a 3B parameter language model designed to push the boundaries of small models. It supports dual mode reasoning, 6 languages and long context. SmolLM3 is a fully open model that offers strong performance at the 3B–4B scale.
1,247 Pulls 5 Tags Updated 9 months ago
a small and good AI model
307 Pulls 1 Tag Updated 6 months ago
SmolLM3 is a 3B parameter language model designed to push the boundaries of small models. It supports 6 languages, advanced reasoning and long context. SmolLM3 is a fully open model that offers strong performance at the 3B–4B scale.
87.7K Pulls 1 Tag Updated 1 year ago
This model is trained to be impolite when speaking in Chinese or sometime English
22 Pulls 1 Tag Updated 9 months ago
495 Pulls 1 Tag Updated 1 year ago
46 Pulls 1 Tag Updated 3 months ago
Extensively pre-trained and instruction fine-tuned version of SmolLM2 360M, now supercharged with German language capabilities!
539 Pulls 4 Tags Updated 1 year ago
Thanks to hugging face for SmolLM 360m while 100m is WIP.
10 Pulls 1 Tag Updated 1 month ago
Huggingface dRAGon Model (https://huggingface.co/llmware/dragon-llama-3.1)
24 Pulls 1 Tag Updated 1 year ago
57 Pulls 1 Tag Updated 10 months ago
SmolLM-135M-GGUF quantized to Q4_0 GGUF for efficient inference.
71 Pulls 1 Tag Updated 6 months ago
21 Pulls 1 Tag Updated 11 months ago
CyberAgentLM3 is a decoder-only language model pre-trained on 2.0 trillion tokens from scratch. CyberAgentLM3-Chat is a fine-tuned model specialized for dialogue use cases.
1,099 Pulls 14 Tags Updated 1 year ago