OLMo 2 is a new family of 7B and 13B models trained on up to 5T tokens. These models are on par with or better than equivalently sized fully open models, and competitive with open-weight models such as Llama 3.1 on English academic benchmarks.
3.7M Pulls 9 Tags Updated 1 year ago
Base model
7 Pulls 3 Tags Updated 1 year ago
A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
3.8M Pulls 5 Tags Updated 1 year ago
Quantized 4-bit version of the original Chocolatine LLM, best performing 13B model on the OpenLLM Leaderboard.
326 Pulls 1 Tag Updated 1 year ago
Single file version with (Dynamic Quants) A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
115 Pulls 4 Tags Updated 1 year ago
7,196 Pulls 2 Tags Updated 1 year ago
4,040 Pulls 5 Tags Updated 1 year ago
(Unsloth Dynamic Quants) A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.
2,063 Pulls 3 Tags Updated 1 year ago
25 Pulls 1 Tag Updated 1 year ago