Q4_K_M and BF16 quantizations of NVIDIA Nemotron Nano 9B v2, NVIDIA’s open 9B reasoning model. Q4 quantized locally from the BF16 source and tuned for a single 16 GB GPU card with maximum KV cache.
227 Pulls 3 Tags Updated 1 month ago
1,384 Pulls 1 Tag Updated 9 months ago