Q4_K_M and BF16 quantizations of NVIDIA Nemotron Nano 9B v2, NVIDIA’s open 9B reasoning model. Q4 quantized locally from the BF16 source and tuned for a single 16 GB GPU card with maximum KV cache.
128 Pulls 3 Tags Updated 3 weeks ago
1,370 Pulls 1 Tag Updated 8 months ago