6 1 week ago

pulled from HF as BF16, requant to IQ4_NL, with combined_en_huge as prompt (English + Coding focus)

tools thinking
ollama run dlasher/openNemo-Cascade-2-30B-A3B:IQ4_NL

Applications

Claude Code
Claude Code ollama launch claude --model dlasher/openNemo-Cascade-2-30B-A3B:IQ4_NL
OpenCode
OpenCode ollama launch opencode --model dlasher/openNemo-Cascade-2-30B-A3B:IQ4_NL
Hermes Agent
Hermes Agent ollama launch hermes --model dlasher/openNemo-Cascade-2-30B-A3B:IQ4_NL
OpenClaw
OpenClaw ollama launch openclaw --model dlasher/openNemo-Cascade-2-30B-A3B:IQ4_NL

Models

View all →

Readme

SOURCE: https://huggingface.co/empero-ai/openNemo-Cascade-2-30B-A3B

openNemo.jpg

Pure-PyTorch drop-in replacement for NVIDIA’s Nemotron-Cascade-2-30B-A3B.

Removes all external CUDA kernel dependencies (mamba-ssm, causal-conv1d) and replaces them with native PyTorch operations, making the model fully compatible with bitsandbytes 4-bit quantization and QLoRA fine-tuning on consumer GPUs.

30B total parameters, 3B active per token. Loads in 17 GB VRAM with 4-bit quantization.