2 days ago

Reduced-refusal Olmo 3 7B Think via Heretic: 48/100 refusals (from 100), KL 0.026. Fully open reasoning model.

ollama run richardyoung/olmo-3-7b-think-heretic

Details

2 days ago

88c7c4da43d0 · 4.5GB

olmo2
·
7.3B
·
Q4_K_M

Readme

Olmo-3-7B-Think-Heretic

Reduced-refusal Olmo 3 7B Think via Heretic: 48⁄100 refusals (from 100), KL 0.026. Fully open reasoning model.

🚀 Overview

This is an abliterated build of allenai/Olmo-3-7B-Think, Ai2’s fully open 7B reasoning model (open weights, data and training code) that thinks before answering. Refusal behavior was reduced using the Heretic library with KL-targeted parameters that preserve the model’s coherence.

Note on refusals: this build still refuses 48 of 100 harmful prompts, more than our other Heretic releases. That is deliberate: the trials that refused less also degraded the model’s normal behavior (KL divergence above 0.10), and we chose to keep the original model’s capabilities intact rather than trade them for fewer refusals.

📊 Abliteration Results

Metric Before After
Refusals 100⁄100 48⁄100
Reduction – 52%
KL Divergence – 0.026

The very low KL divergence (0.026, far below the 0.5 “damage” threshold) means the model retains essentially all of its original capabilities and coherence.

🎯 Key Features

  • Reduced censorship: 52% fewer refusals on typical “unsafe” prompts
  • Near-zero quality loss: KL 0.026
  • Fully open model: weights, data and code
  • Reasoning model with a thinking phase
  • 64K-token context
  • Apache-2.0 license
  • Reproducible: full reproduction data on Hugging Face

🏷️ Available Versions

Tag Size BPW Notes
latest / Q4_K_M 4.5 GB 4.85 Recommended
Q5_K_M 5.2 GB 5.68 Higher quality
Q6_K 6.0 GB 6.56 Very high quality
Q8_0 7.8 GB 8.5 Near-lossless

💻 Quick Start

ollama run richardyoung/olmo-3-7b-think-heretic           # recommended (Q4_K_M)
ollama run richardyoung/olmo-3-7b-think-heretic:Q8_0      # near-lossless

🛠️ Use Cases

  • Research on reasoning and alignment in a fully open model
  • Math and logic with visible reasoning
  • Red-teaming

📋 System Requirements

VRAM Recommended tier
6 GB+ Q4_K_M (4.5 GB file)
9 GB+ Q8_0 (7.8 GB file)

🔧 Technical Details

  • Base Model: allenai/Olmo-3-7B-Think
  • Parameters: 7.3B (olmo3 architecture)
  • Context Length: 64K tokens
  • Quantization: GGUF via llama.cpp (text generation)
  • Abliteration: Heretic v2.0.0.dev0 by p-e-w (Trial 136: 100→48 refusals @ KL 0.026)
  • License: Apache-2.0 (inherited from the base model)

⚠️ Disclaimer

This model has reduced safety guardrails. The removal of refusal behavior means it will engage with a wider range of prompts. Use responsibly and in accordance with applicable laws and regulations.

🙏 Acknowledgments

  • Base Model: Allen Institute for AI (Ai2)
  • Abliteration: Heretic by p-e-w
  • Quantization: llama.cpp

Built & maintained by Richard Young · DeepNeuro