39 2 weeks ago

Gemma 4 31B (Q4_0) optimized for Hermes Agent — 64K context, 8192 max tokens, flash attention + Q8 KV cache. Fits 24 GB VRAM with optimizations.

vision tools thinking
91d0bf03677b · 78B
{
"num_ctx": 65536,
"num_predict": 8192,
"temperature": 0.7,
"top_k": 64,
"top_p": 0.9
}