1 1 month ago

ollama run cmsmanhattan/jirack-ultra-32b-q4

Models

View all →

Readme

JiRack Ultra 32B Q4_K_M

Large ~32B model from CMS Manhattan. BitNet-style ternary path, extended tokenizer with Routing, Tool-call and Robotics tags. Production Q4_K_M build for practical multi-GPU and high-RAM inference.

HF: https://huggingface.co/CMSManhattan/JiRackUltra_32b

Run

ollama run cmsmanhattan/jirack-ultra-32b-q4

What it is

  • Flagship mid/large size in the Ultra line — stronger reasoning, coding, and tool use than 14B
  • Same JiRack tokenizer tags: routing, tools, robotics
  • Q4_K_M quant for usable quality at lower memory than FP16
  • Fits serious RAG, multi-agent, and Spring / Java tool-call stacks
  • American ternary stack from CMS Manhattan (New York)

Hardware

Q4_K_M: roughly 18–20 GB on disk.
Recommend 24–32 GB VRAM (GPU) or 48+ GB system RAM (CPU).

Contact

grabko@cmsmanhattan.com
+1 (516) 777-0945
New York, USA

License

MIT on weights unless stated otherwise on the HF card. Commercial terms for UI / enterprise on request.