463 3 weeks ago

ollama run n0404n0404/qwen3.6-finetune-qwen3.8-8287f4-heretic-7aaada:q2_k

Details

3 weeks ago

e3f5beb42fd3 · 11GB ·

qwen35
·
26.9B
·
Q2_K

Readme

Qwen3.6-27B Qwen3.8-Distill — Heretic (Abliterated)

Text-only GGUF build of a Qwen3.8-Max distillation finetune of Qwen/Qwen3.6-27B, decensored with Heretic (automatic abliteration via directional ablation with per-model parameter optimization).

The finetune is a LoRA trained on r0b0tlab/qwen3.8-max-distillation-50k — 50k reasoning traces distilled from Qwen 3.8 Max — merged into the full weights with PEFT before abliteration. The goal is to transfer Qwen 3.8’s reasoning style onto the 27B base. This build removes the refusal direction on top of that — the model answers questions the original would decline. Reasoning (<think>) behavior is preserved and parsed natively by Ollama.

Tags

Tag Quantization File size (approx.) Suggested VRAM
q2_k 2-bit K-quant ~10.7 GB 16 GB
q3_k_m 3-bit K-quant ~13.3 GB 16 GB
Q4_K_M 4-bit K-quant (recommended) ~16.5 GB 24 GB
q6_k 6-bit K-quant (near-lossless) ~22.1 GB 32 GB
ollama run n0404n0404/qwen3.6-finetune-qwen3.8-8287f4-heretic-7aaada:Q4_K_M

Long context is unusually cheap on this architecture: 48 of the 64 layers use linear attention, so the KV cache only costs ~64 KiB per token (about 2 GiB at 32k context). A 32 GB GPU (e.g. RTX 5090) runs q6_k fully offloaded at 32k–64k context; 24 GB cards are comfortable with Q4_K_M.

Model details

  • Base: Qwen/Qwen3.6-27B → Qwen3.8-Max distillation LoRA (merged with PEFT) → Heretic abliteration
  • Training data: r0b0tlab/qwen3.8-max-distillation-50k
  • Architecture: qwen35 hybrid — 64 layers (48 linear attention + 16 full attention), GQA with 4 KV heads × 256 head dim, up to 262,144 context
  • Thinking: reasoning model; served with Ollama’s native qwen3.5 renderer/parser, so thinking content is separated from the final answer
  • Requires: Ollama 0.31.2 or newer

What was changed vs. the original

  1. Distillation finetune — a LoRA trained on 50k Qwen 3.8 Max reasoning traces was merged into the base weights (PEFT merge_and_unload).
  2. Abliteration — refusal behavior removed with Heretic (run 7aaada): refusals on a 100-prompt harmful set dropped from 94100 to 5100, with a KL divergence of only 0.0056 vs. the pre-abliteration model — output distribution on normal prompts is essentially unchanged.
  3. Text-only — the original is multimodal (vision-language). The vision tower is not included in this GGUF; image input is not supported.
  4. No MTP head — the multi-token-prediction (speculative decoding) head was dropped during conversion (--no-mtp); it is unused by Ollama.

Converted with llama.cpp convert_hf_to_gguf.py + llama-quantize.

Sanity check after abliteration + quantization: GSM8K (300-sample subset, LM Evaluation Harness via Ollama’s OpenAI-compatible API, max_gen_toks 8192) scores 97.33% strict exact-match — on par with the published base-model level, i.e. no measurable capability loss from the finetune, abliteration, or Q4 quantization.

Responsible use

Safety guardrails have been intentionally weakened. This model will comply with requests the original model would refuse. You are responsible for how you use it and for complying with applicable laws and the upstream license. Not recommended for user-facing deployments without your own moderation layer.

License & credits


繁體中文說明

Qwen/Qwen3.6-27BQwen3.8-Max 蒸餾微調 + 去審查(abliterated)版本,使用 Heretic 自動化處理,轉為純文字 GGUF 供 Ollama 使用。

  • 微調資料集:r0b0tlab/qwen3.8-max-distillation-50k(自 Qwen 3.8 Max 蒸餾的 5 萬筆推理軌跡),以 LoRA 訓練後用 PEFT 合併為完整權重
  • 本版本移除了拒答行為(Heretic run 7aaada):拒答率 941005100,KL 散度僅 0.0056,一般提示的輸出分布幾乎不變
  • <think> 推理功能完整保留,由 Ollama 原生解析
  • 混合線性注意力架構,KV cache 僅約 64 KiB/token,長上下文非常省 VRAM
  • 建議:16 GB 顯卡用 q2_k/q3_k_m,24 GB 用 Q4_K_M,32 GB(如 RTX 5090)可用 q6_k 開 32k–64k 上下文
  • 量化後實測 GSM8K(300 題抽樣)strict 97.33%,與原版基準持平——微調、消融、Q4 量化皆無能力損失
  • 不含視覺功能(原模型為多模態,本 GGUF 僅文字)
  • 需要 Ollama 0.31.2 以上版本

負責任使用:本模型已刻意弱化安全防護,會回應原模型拒絕的請求。使用者須自行承擔使用責任並遵守相關法律與授權條款。不建議在未加自有審核層的情況下對外部署。