1 1 month ago

ollama run cmsmanhattan/jirack-ultra-32b-q4

Details

1 month ago

f227779dc7c2 · 20GB

qwen2
·
32.8B
·
Q4_K_M
<|User|>{{ .Prompt }}<|Assistant|>
You are JiRack, a helpful and honest AI assistant.
{ "num_ctx": 8192, "repeat_penalty": 1.1, "stop": [ "<|end▁of▁sentence|>

Readme

JiRack Ultra 32B Q4_K_M

Large ~32B model from CMS Manhattan. BitNet-style ternary path, extended tokenizer with Routing, Tool-call and Robotics tags. Production Q4_K_M build for practical multi-GPU and high-RAM inference.

HF: https://huggingface.co/CMSManhattan/JiRackUltra_32b

Run

ollama run cmsmanhattan/jirack-ultra-32b-q4

What it is

  • Flagship mid/large size in the Ultra line — stronger reasoning, coding, and tool use than 14B
  • Same JiRack tokenizer tags: routing, tools, robotics
  • Q4_K_M quant for usable quality at lower memory than FP16
  • Fits serious RAG, multi-agent, and Spring / Java tool-call stacks
  • American ternary stack from CMS Manhattan (New York)

Hardware

Q4_K_M: roughly 18–20 GB on disk.
Recommend 24–32 GB VRAM (GPU) or 48+ GB system RAM (CPU).

Contact

grabko@cmsmanhattan.com
+1 (516) 777-0945
New York, USA

License

MIT on weights unless stated otherwise on the HF card. Commercial terms for UI / enterprise on request.