19 2 days ago

Ornith-1.5-9B quantized to IQ2_M (3.77 GB) — the most aggressive compression available for this model. Ideal for low-VRAM GPUs. Runs with a 16k context window while leaving 2.5 GB of VRAM free. Slightly faster at the cost of some precision

vision
c48d2a5f1396 · 60B
{
"num_ctx": 16384,
"temperature": 0.5,
"top_k": 20,
"top_p": 0.95
}