7 yesterday

Uncensored LFM2.5-1.2B-Instruct via Heretic: 3/100 refusals (from 98), KL 0.059. Tiny and fast; runs on almost anything.

ollama run richardyoung/lfm2.5-1.2b-instruct-heretic:Q6_K

Details

yesterday

c626e9f4b692 ยท 963MB

lfm2
ยท
1.17B
ยท
Q6_K

Readme

LFM2.5-1.2B-Instruct-Heretic

Uncensored LFM2.5-1.2B-Instruct via Heretic: 3โ„100 refusals (from 98), KL 0.059. Tiny and fast; runs on almost anything.

๐Ÿš€ Overview

This is an abliterated build of LiquidAI/LFM2.5-1.2B-Instruct, Liquid AIโ€™s 1.2B edge model with a hybrid convolution + attention architecture, built for fast on-device inference. Refusal behavior was reduced using the Heretic library with KL-targeted parameters that preserve the modelโ€™s coherence.

๐Ÿ“Š Abliteration Results

Metric Before After
Refusals 98โ„100 3โ„100
Reduction โ€“ 97%
KL Divergence โ€“ 0.059

The very low KL divergence (0.059, far below the 0.5 โ€œdamageโ€ threshold) means the model retains essentially all of its original capabilities and coherence.

๐ŸŽฏ Key Features

  • Reduced censorship: 97% fewer refusals on typical โ€œunsafeโ€ prompts
  • Near-zero quality loss: KL 0.059
  • Tiny footprint: runs comfortably on CPU or any GPU
  • 128K-token context
  • Fast hybrid conv + attention architecture
  • Reproducible: full reproduction data on Hugging Face

๐Ÿท๏ธ Available Versions

Tag Size BPW Notes
latest / Q4_K_M 0.7 GB 4.85 Recommended
Q5_K_M 0.8 GB 5.68 Higher quality
Q6_K 1.0 GB 6.56 Very high quality
Q8_0 1.2 GB 8.5 Near-lossless

๐Ÿ’ป Quick Start

ollama run richardyoung/lfm2.5-1.2b-instruct-heretic           # recommended (Q4_K_M)
ollama run richardyoung/lfm2.5-1.2b-instruct-heretic:Q8_0      # near-lossless

๐Ÿ› ๏ธ Use Cases

  • On-device and edge assistants without stock refusals
  • Research and red-teaming on small models
  • Creative writing

๐Ÿ“‹ System Requirements

VRAM Recommended tier
4 GB+ Q4_K_M (0.7 GB file)
4 GB+ Q8_0 (1.2 GB file)

๐Ÿ”ง Technical Details

  • Base Model: LiquidAI/LFM2.5-1.2B-Instruct
  • Parameters: 1.2B (lfm2 architecture)
  • Context Length: 125K tokens
  • Quantization: GGUF via llama.cpp (text generation)
  • Abliteration: Heretic v2.0.0.dev0 by p-e-w (Trial 78: 98โ†’3 refusals @ KL 0.059)
  • License: LFM Open License v1.0 (inherited from the base model)

โš ๏ธ Disclaimer

This model has reduced safety guardrails. The removal of refusal behavior means it will engage with a wider range of prompts. Use responsibly and in accordance with applicable laws and regulations.

๐Ÿ™ Acknowledgments

  • Base Model: Liquid AI
  • Abliteration: Heretic by p-e-w
  • Quantization: llama.cpp

Built & maintained by Richard Young ยท DeepNeuro