82 2 weeks ago

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities.

ollama run AntLing/Ling-3.0-flash:Q4_K_M

Details

2 weeks ago

7f4de67dae5e · 77GB ·

bailingmoe3
·
127B
·
Q4_K_M
{ "temperature": 0.6, "top_k": 20, "top_p": 0.95 }

Readme

No readme