6 1 week ago

pulled from HF as BF16, requant to IQ4_NL, with combined_en_huge as prompt (English + Coding focus)

tools thinking
ollama run dlasher/openNemo-Cascade-2-30B-A3B:IQ4_NL

Details

1 week ago

d84f9a6aa79e · 18GB

nemotron_h_moe
·
31.6B
·
IQ4_NL
{ "min_p": 0.05, "num_batch": 2048, "num_ctx": 262144, "presence_penalty": 0, "r
{{- if .System }}<|im_start|>system {{ .System }}<|im_end|> {{- end }} {{- range .Messages }} {{- if

Readme

SOURCE: https://huggingface.co/empero-ai/openNemo-Cascade-2-30B-A3B

openNemo.jpg

Pure-PyTorch drop-in replacement for NVIDIA’s Nemotron-Cascade-2-30B-A3B.

Removes all external CUDA kernel dependencies (mamba-ssm, causal-conv1d) and replaces them with native PyTorch operations, making the model fully compatible with bitsandbytes 4-bit quantization and QLoRA fine-tuning on consumer GPUs.

30B total parameters, 3B active per token. Loads in 17 GB VRAM with 4-bit quantization.