44 2 months ago

Thinking model - efficient

ollama run Yarflam/SmolLM3-Q4KM

Details

2 months ago

708d818fbbcd · 1.9GB ·

smollm3
·
3.08B
·
Q4_K_M
{ "num_ctx": 16384, "stop": [ "<|im_end|>", "<|im_start|>" ], "tempe
{{ if .System }}<|im_start|>system {{ .System }}<|im_end|> {{ end }}{{ range .Messages }}<|im_start|

Readme

Original model from https://huggingface.co/ggml-org/SmolLM3-3B-GGUF.

Thinking model – efficient when you don’t have enough GPU on your computer. 🤪