24 Downloads Updated yesterday
ollama run richardyoung/muse-glimmer-30b-heretic:Q4_K_M
Reduced-refusal Muse Glimmer 30B via Heretic: 52⁄100 refusals (from 99), KL 0.063. Vision, Apache-2.0.
An abliterated build of Meta’s Muse Glimmer 30B, a 29.6B dense vision-language model built for local agents: it reads images (screenshots, charts, documents), reasons before it answers with adjustable effort, and calls tools. The refusal behavior was reduced with a reproducible 200-trial Heretic run; the vision encoder is untouched, and every tag here includes the vision projector.
Note on refusals: this build still refuses 52 of 100 harmful prompts, more than our other Heretic releases. That is deliberate: the trials that refused less also degraded the model’s normal behavior (KL divergence above 0.10), and we chose to keep the original model’s capabilities intact rather than trade them for fewer refusals.
Refusals are scored on the model’s actual answer, with its reasoning block skipped (the approach Heretic uses for gpt-oss). Scoring the start of the reasoning instead counts far fewer refusals even for the unmodified model, so numbers from runs that did that aren’t directly comparable.
| Metric | Value |
|---|---|
| Refusals (original) | 99⁄100 |
| Refusals (this model) | 52⁄100 |
| Reduction | 47% |
| KL divergence | 0.063 |
The very low KL divergence (0.063, far below the 0.5 “damage” threshold) means the model retains essentially all of its original capabilities and coherence.
think to low, medium, high, max or falsenum_ctx if you have the memory)| Tag | Size | BPW | Notes |
|---|---|---|---|
IQ3_M |
16 GB | 3.68 | Smallest |
IQ4_XS |
19 GB | 4.37 | Great quality-size balance |
Q4_K_M / latest |
20 GB | 4.86 | Recommended |
Q5_K_M |
23 GB | 5.69 | Higher quality |
Q6_K |
26 GB | 6.56 | Very high quality |
Q8_0 |
33 GB | 8.50 | Near-lossless |
Sizes include the 3.8 GB vision projector.
Requires Ollama 0.32.8 or newer.
ollama run richardyoung/muse-glimmer-30b-heretic
With an image: drag it into the chat, or put the path in the prompt:
ollama run richardyoung/muse-glimmer-30b-heretic "Transcribe this receipt: ./receipt.png"
API, with less thinking:
curl http://localhost:11434/api/chat -d '{
"model": "richardyoung/muse-glimmer-30b-heretic",
"think": "low",
"messages": [{"role": "user", "content": "Explain abliteration in simple terms."}]
}'
| VRAM | Suggested tag |
|---|---|
| 16 GB | IQ3_M (text-only, or with partial CPU offload for images) |
| 24 GB | IQ4_XS or Q4_K_M with vision (Q4_K_M uses ~20 GB at 16K context) |
| 32 GB | Q5_K_M or Q6_K |
| 40 GB+ | Q8_0 |
Apple Silicon: similar amounts of unified memory. CPU-only works but is slow for a 30B dense model.
This model has reduced safety guardrails. The removal of refusal behavior means it will engage with a wider range of prompts. Use responsibly and in accordance with applicable laws and regulations.
The original model ships with Meta’s Muse Glimmer Usage Policy. The weights are licensed Apache-2.0; you are responsible for making sure your use complies with the licence, that policy, and applicable laws.
Base Model: Meta · Abliteration: Heretic by p-e-w · Quantization: llama.cpp
Built & maintained by Richard Young · DeepNeuro