4 1 month ago

ollama run cmsmanhattan/JiRackDeltaNet_27b-q4-reasoning

Models

View all →

Readme

JiRackDeltaNet 27B Reasoning Local agentic / tool-calling model from CMS Manhattan. Ternary-oriented build of Qwen/Qwen3.8-27B.
Named DeltaNet after the Qwen3.8 Gated DeltaNet stack (Gated DeltaNet + gated attention).
Built to keep tool calling and reasoning on hardware you control — so a vendor model change cannot silently break finance, ops, or audit workflows. 27B parameters Native context: 256K Reasoning / thinking model Variants: Q4, Q3, Q2 for different VRAM budgets

More models: huggingface.co/CMSManhattan

Run

ollama run cmsmanhattan/JiRackDeltaNet_27b-q4-reasoning

Q3 (smaller):

ollama run cmsmanhattan/JiRackDeltaNet_27b-q3-reasoning

Q2 (smallest):

ollama run cmsmanhattan/JiRackDeltaNet_27b-q2-reasoning

API

curl http://localhost:11434/api/chat -d '{
  "model": "cmsmanhattan/JiRackDeltaNet_27b-q4-reasoning",
  "messages": [{"role": "user", "content": "Plan a 3-step tool-using workflow to reconcile two ledgers."}],
  "think": true
}'

Python

from ollama import chat

response = chat(
    model="cmsmanhattan/JiRackDeltaNet_27b-q4-reasoning",
    messages=[{"role": "user", "content": "Explain why local tool calling matters in finance agents."}],
    think=True,
)
print(response.message.content)

Why this model Hosted APIs can change or retire a model overnight. If your agent posts, closes, or certifies money, a broken tool call is not a chatbot mistake — it is books, controls, and audit trail. JiRackDeltaNet is for teams that want: local inference stable tool calling reasoning traces you can keep on-prem

Treat cloud models as a prototype surface. Own the inference for anything that moves or certifies money.

Variants Model Link Use Q4 reasoning https://ollama.com/cmsmanhattan/JiRackDeltaNet_27b-q4-reasoning Best quality / ~17GB Q3 reasoning https://ollama.com/cmsmanhattan/JiRackDeltaNet_27b-q3-reasoning Mid VRAM Q2 reasoning https://ollama.com/cmsmanhattan/JiRackDeltaNet_27b-q2-reasoning Tight VRAM / edge

Q4 page currently lists 17GB · 256K context · text.

Base model Source: Qwen/Qwen3.8-27B Architecture: Gated DeltaNet + gated attention (Qwen3.8) License of the base model: Apache 2.0

This release is a CMS Manhattan local/quantized reasoning build for Ollama. It is not an official Qwen checkpoint.

Publisher CMS Manhattan
https://huggingface.co/CMSManhattan
https://ollama.com/cmsmanhattan