159 Downloads Updated 7 months ago
ollama run richardyoung/gemma-7b-it-abliterated:Q4_K_M
Abliterated (uncensored) version of Gemma-7B-IT for unrestricted conversations and creative writing.
This represents an abliterated variant of google/gemma-7b-it, featuring significantly diminished refusal mechanisms through targeted weight adjustment. The abliteration process uses the Heretic library with conservative parameters to preserve model coherence while removing restrictions.
| Metric | Before | After |
|---|---|---|
| Refusals | TBD | TBD |
| Reduction | , | TBD |
| KL Divergence | , | TBD |
Refusal metrics pending re-measurement.
The goal of abliteration is to minimize KL divergence (keeping it low) so the model maintains its original capabilities and coherence while reducing refusals.
| Tag | Size | BPW | Notes |
|---|---|---|---|
latest / Q4_K_M |
5.3 GB | 4.85 | Recommended |
BPW reference guide (bits per weight, for choosing a quantization):
| Quant | BPW | Notes |
|---|---|---|
| IQ3_M | 3.66 | Smallest, for low VRAM |
| IQ4_XS | 4.25 | Great quality/size balance |
| Q4_K_M | 4.85 | Recommended balance |
| Q5_K_M | 5.68 | Higher quality |
| Q6_K | 6.56 | Very high quality |
| Q8_0 | 8.5 | Near-lossless |
ollama run richardyoung/gemma-7b-it-abliterated
ollama run richardyoung/gemma-7b-it-abliterated:latest
| VRAM | Performance |
|---|---|
| 6 GB | Slow, may swap to CPU |
| 8 GB | Good performance |
| 12 GB+ | Excellent performance |
Base Model: google/gemma-7b-it | Parameters: ~8.5B (7B class) | Context: 8,192 tokens | Quantization: Q4_K_M (4.85 bits per weight)
This model has reduced safety guardrails. The removal of refusal behavior means the model will engage with a wider range of prompts. Use responsibly and in accordance with applicable laws and regulations. Use is subject to the Gemma license.
Built & maintained by Richard Young ยท DeepNeuro