1,531 1 month ago

Decensored Qwen2.5-3B-Instruct with 2/100 refusals via Heretic abliteration. General-purpose 3B model for local use.

tools
ollama run R4C3R/qwen2.5-3b-heretic

Applications

Claude Code
Claude Code ollama launch claude --model R4C3R/qwen2.5-3b-heretic
OpenCode
OpenCode ollama launch opencode --model R4C3R/qwen2.5-3b-heretic
Hermes Agent
Hermes Agent ollama launch hermes --model R4C3R/qwen2.5-3b-heretic
OpenClaw
OpenClaw ollama launch openclaw --model R4C3R/qwen2.5-3b-heretic

Models

View all →

Readme

🏴 Qwen2.5-3B-Heretic

Heretic Series by RACER IS OP


Qwen2.5-3B-Heretic is a decensored variant of Qwen’s popular 3B instruct model. Refusals dropped from 96100 to 2100 with minimal capability loss (KL 0.13) — one of the most effective ablations in the series. Runs comfortably on CPU with Q4 quantization.

Who this is for — developers who want a compact 3B general-purpose model that answers directly instead of refusing. Great for local agents, roleplay, edge deployment, or anything blocked by RLHF over-refusal.


Quickstart

ollama run R4C3R/qwen2.5-3b-heretic

VRAM Guide (choose your quant)

Quant VRAM When to use
q4_k_m ~1.8 GB Best balance — runs on any system
q5_k_m ~2.1 GB Higher quality, still lightweight
q8_0 ~3.2 GB Maximum quality, CPU-friendly

Have a 4GB GPU? Start with q4_k_m. On CPU only? q8_0 gives you full quality.

Links

Hugging Face · sdad.pro · Heretic

License: qwen-research


Made with love by RACER IS OP — follow for more uncensored models in the Heretic Series.