37 6 days ago

Q4_K_M and BF16 quantizations of NVIDIA Nemotron Nano 9B v2, NVIDIA’s open 9B reasoning model. Q4 quantized locally from the BF16 source and tuned for a single 16 GB GPU card with maximum KV cache.

tools thinking