741 1 month ago

vision tools thinking
ollama run jacokon/qwen3.8-27b-heretic-ara

Details

1 month ago

c0a5bbd6966e · 18GB

qwen35
·
27.3B
·
Q4_K_M
clip
·
461M
·
F16
{{- if .Messages }} {{- if or .System .Tools }}<|im_start|>system {{- if .System }} {{ .System }} {{
{ "num_ctx": 65536, "repeat_penalty": 1, "stop": [ "<|im_end|>", "<|endo

Readme

Qwen3.8-27B Heretic (Vision + Coding Agent Optimized)

A decensored/abliterated multimodal vision-language model (VLM) based on Qwen3.8-27B, fine-tuned and configured for seamless execution in autonomous coding agents (such as Claude Code, OpenCode, Aider, and Hermes) with full Multimodal Vision support.


📦 Available Tags & Quantizations

Tag Quantization & Tech Size Recommended Hardware / Target Use Case
latest (or q4_k_m) Q4_K_M + mmproj-f16 17 GB 24 GB ~ 32 GB+ VRAM (RTX 3090, 4090, 5090, Apple Silicon 36GB+). Peak precision, maximum tool-calling and code generation accuracy.
iq3_m ⭐ (Recommended for 16GB) IQ3_M (I-Matrix) + MTP + mmproj-Q8_0 13 GB 16 GB VRAM (RTX 4080, 4070 Ti Super, 4060 Ti 16GB, RX 7800 XT). Uses Importance Matrix (I-Matrix) for near-Q4 quality + Multi-Token Prediction (MTP) for smaller VRAM usage & faster generation.
q3_k_m Q3_K_M + mmproj-f16 14 GB 16 GB VRAM standard K-quantization fallback.

⚡ Highlights & Key Features

This release features custom, robust Jinja template patches and native Ollama layer packaging designed for real-world agentic workflows:

  1. Importance Matrix & MTP Support (iq3_m):
    • Quantized using I-Matrix calibration, preserving critical attention weights in code syntax, math, and JSON schema.
    • Multi-Token Prediction (MTP) enabled for fast speculative inference and smaller memory footprint (~12.7 GB total with vision).
  2. Integrated Multimodal Vision (mmproj):
    • Packaged with vision projector weights. Supports screenshot inspection, UI debugging, diagram analysis, and image inputs directly within Claude Code (/image) and Agent workflows.
  3. Deterministic Agent Tool Calling:
    • Uses clean JSON-based tool call encapsulation (<tool_call>{"name": ..., "arguments": ...}</tool_call>), preventing parsing anomalies, JSON decode errors, and streaming truncation across Anthropic /v1/messages and OpenAI /v1/chat/completions compatibility layers.
  4. Robust Multi-Turn Template (Zero Assertion Crashes):
    • Stripped fragile Jinja assertions (raise_exception), supporting middle-turn system reminders, multi-step tool turns without immediate user queries, unescaped regex/log tool outputs, and Anthropic reasoning effort parameters (high, medium, low).
  5. Context Window:
    • Default 64k (num_ctx 65536), dynamically adjustable up to 256k (262,144 tokens) based on available hardware.

🚀 Quick Start

1. Launch with Claude Code (Coding & Vision)

Recommended for 16GB VRAM (IQ3_M + MTP):

ollama launch claude --model jacokon/qwen3.8-27b-heretic-ara:iq3_m

Standard Version (24GB+ VRAM):

ollama launch claude --model jacokon/qwen3.8-27b-heretic-ara:latest

2. Launch with OpenCode

# 16GB VRAM (IQ3_M)
ollama launch opencode --model jacokon/qwen3.8-27b-heretic-ara:iq3_m

# 24GB+ VRAM (Q4_K_M)
ollama launch opencode --model jacokon/qwen3.8-27b-heretic-ara:latest

3. Run in Terminal

ollama run jacokon/qwen3.8-27b-heretic-ara:iq3_m

4. Multimodal Vision via Python (OpenAI / Anthropic API)

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"
)

# Example: Analyze a screenshot or image
response = client.chat.completions.create(
    model="jacokon/qwen3.8-27b-heretic-ara:iq3_m",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Analyze this screenshot and explain what is on the screen:"},
                {
                    "type": "image_url",
                    "image_url": {"url": "data:image/png;base64,<BASE64_IMAGE_DATA>"}
                }
            ]
        }
    ],
    temperature=0.6,
)

print(response.choices[0].message.content)

💻 Hardware Guidelines

  • jacokon/qwen3.8-27b-heretic-ara:iq3_m (IQ3_M + MTP): ⭐ Best choice for 16 GB VRAM GPUs (e.g., RTX 4080, RTX 4070 Ti Super, RTX 4060 Ti 16GB, RX 7800 XT). Total footprint is only ~12.1 GB (model) + ~0.59 GB (vision), leaving ~3.3 GB of VRAM headroom for 64k Context KV Cache and OS desktop display, achieving 100% full GPU acceleration with zero CPU offloading.
  • jacokon/qwen3.8-27b-heretic-ara:latest (Q4_K_M): Ideal for 24 GB+ VRAM GPUs (RTX 3090, RTX 4090, RTX 5090, or Apple Silicon with 36GB+ unified memory) to achieve full precision and maximum coding intelligence.
  • jacokon/qwen3.8-27b-heretic-ara:q3_k_m (Q3_K_M): Alternative standard 3-bit K-quantization for 16 GB setups.

📜 Credits & Acknowledgments