82 2 weeks ago

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities.

ollama run AntLing/Ling-3.0-flash:Q6_K

Details

2 weeks ago

cac9f3dc6ddd · 105GB ·

bailingmoe3
·
127B
·
Q6_K
{ "temperature": 0.6, "top_k": 20, "top_p": 0.95 }

Readme

No readme