24 yesterday

Reduced-refusal Muse Glimmer 30B via Heretic: 52/100 refusals (from 99), KL 0.063. Vision, Apache-2.0.

vision tools thinking
ollama run richardyoung/muse-glimmer-30b-heretic:Q5_K_M

Details

yesterday

bd4deead63b4 · 24GB

muse-glimmer
·
27.9B
·
Q5_K_M
clip
·
1.92B
·
F16
{ "num_ctx": 16384, "temperature": 1, "top_k": 64, "top_p": 0.95 }

Readme

Muse Glimmer 30B Heretic

Reduced-refusal Muse Glimmer 30B via Heretic: 52⁄100 refusals (from 99), KL 0.063. Vision, Apache-2.0.

🚀 Overview

An abliterated build of Meta’s Muse Glimmer 30B, a 29.6B dense vision-language model built for local agents: it reads images (screenshots, charts, documents), reasons before it answers with adjustable effort, and calls tools. The refusal behavior was reduced with a reproducible 200-trial Heretic run; the vision encoder is untouched, and every tag here includes the vision projector.

Note on refusals: this build still refuses 52 of 100 harmful prompts, more than our other Heretic releases. That is deliberate: the trials that refused less also degraded the model’s normal behavior (KL divergence above 0.10), and we chose to keep the original model’s capabilities intact rather than trade them for fewer refusals.

Refusals are scored on the model’s actual answer, with its reasoning block skipped (the approach Heretic uses for gpt-oss). Scoring the start of the reasoning instead counts far fewer refusals even for the unmodified model, so numbers from runs that did that aren’t directly comparable.

📊 Abliteration Results

Metric Value
Refusals (original) 99⁄100
Refusals (this model) 52⁄100
Reduction 47%
KL divergence 0.063

The very low KL divergence (0.063, far below the 0.5 “damage” threshold) means the model retains essentially all of its original capabilities and coherence.

🎯 Key Features

  • Vision: text + image input through Meta’s ~1.8B perception encoder; tested on document transcription and image description
  • Reasoning with adjustable effort: thinking on by default (high); set think to low, medium, high, max or false
  • Tool calling and agentic workflows (the original is tuned for MCP-style tool use and coding agents)
  • 131K context window (tags default to 16K; raise num_ctx if you have the memory)
  • Multilingual: trained on data from more than 100 languages
  • Apache-2.0, reproducible: settings and seed are published with the full weights

🏷️ Available Versions

Tag Size BPW Notes
IQ3_M 16 GB 3.68 Smallest
IQ4_XS 19 GB 4.37 Great quality-size balance
Q4_K_M / latest 20 GB 4.86 Recommended
Q5_K_M 23 GB 5.69 Higher quality
Q6_K 26 GB 6.56 Very high quality
Q8_0 33 GB 8.50 Near-lossless

Sizes include the 3.8 GB vision projector.

💻 Quick Start

Requires Ollama 0.32.8 or newer.

ollama run richardyoung/muse-glimmer-30b-heretic

With an image: drag it into the chat, or put the path in the prompt:

ollama run richardyoung/muse-glimmer-30b-heretic "Transcribe this receipt: ./receipt.png"

API, with less thinking:

curl http://localhost:11434/api/chat -d '{
  "model": "richardyoung/muse-glimmer-30b-heretic",
  "think": "low",
  "messages": [{"role": "user", "content": "Explain abliteration in simple terms."}]
}'

🛠️ Use Cases

  • Document, screenshot and chart understanding (OCR, invoices, UI screenshots)
  • Local agents and tool calling without cloud access
  • Coding assistance
  • Creative writing and fiction without reflexive refusals
  • Alignment and refusal research

📋 System Requirements

VRAM Suggested tag
16 GB IQ3_M (text-only, or with partial CPU offload for images)
24 GB IQ4_XS or Q4_K_M with vision (Q4_K_M uses ~20 GB at 16K context)
32 GB Q5_K_M or Q6_K
40 GB+ Q8_0

Apple Silicon: similar amounts of unified memory. CPU-only works but is slow for a 30B dense model.

🔧 Technical Details

⚠️ Disclaimer

This model has reduced safety guardrails. The removal of refusal behavior means it will engage with a wider range of prompts. Use responsibly and in accordance with applicable laws and regulations.

The original model ships with Meta’s Muse Glimmer Usage Policy. The weights are licensed Apache-2.0; you are responsible for making sure your use complies with the licence, that policy, and applicable laws.

🙏 Acknowledgments

Base Model: Meta · Abliteration: Heretic by p-e-w · Quantization: llama.cpp

Built & maintained by Richard Young · DeepNeuro