463 Downloads Updated 3 weeks ago
ollama run n0404n0404/qwen3.6-finetune-qwen3.8-8287f4-heretic-7aaada:q2_k
Updated 3 weeks ago
3 weeks ago
e3f5beb42fd3 · 11GB ·
Text-only GGUF build of a Qwen3.8-Max distillation finetune of Qwen/Qwen3.6-27B, decensored with Heretic (automatic abliteration via directional ablation with per-model parameter optimization).
The finetune is a LoRA trained on r0b0tlab/qwen3.8-max-distillation-50k — 50k reasoning traces distilled from Qwen 3.8 Max — merged into the full weights with PEFT before abliteration. The goal is to transfer Qwen 3.8’s reasoning style onto the 27B base. This build removes the refusal direction on top of that — the model answers questions the original would decline. Reasoning (<think>) behavior is preserved and parsed natively by Ollama.
| Tag | Quantization | File size (approx.) | Suggested VRAM |
|---|---|---|---|
q2_k |
2-bit K-quant | ~10.7 GB | 16 GB |
q3_k_m |
3-bit K-quant | ~13.3 GB | 16 GB |
Q4_K_M |
4-bit K-quant (recommended) | ~16.5 GB | 24 GB |
q6_k |
6-bit K-quant (near-lossless) | ~22.1 GB | 32 GB |
ollama run n0404n0404/qwen3.6-finetune-qwen3.8-8287f4-heretic-7aaada:Q4_K_M
Long context is unusually cheap on this architecture: 48 of the 64 layers use linear attention, so the KV cache only costs ~64 KiB per token (about 2 GiB at 32k context). A 32 GB GPU (e.g. RTX 5090) runs q6_k fully offloaded at 32k–64k context; 24 GB cards are comfortable with Q4_K_M.
qwen35 hybrid — 64 layers (48 linear attention + 16 full attention), GQA with 4 KV heads × 256 head dim, up to 262,144 contextqwen3.5 renderer/parser, so thinking content is separated from the final answermerge_and_unload).7aaada): refusals on a 100-prompt harmful set dropped from 94⁄100 to 5⁄100, with a KL divergence of only 0.0056 vs. the pre-abliteration model — output distribution on normal prompts is essentially unchanged.--no-mtp); it is unused by Ollama.Converted with llama.cpp convert_hf_to_gguf.py + llama-quantize.
Sanity check after abliteration + quantization: GSM8K (300-sample subset, LM Evaluation Harness via Ollama’s OpenAI-compatible API, max_gen_toks 8192) scores 97.33% strict exact-match — on par with the published base-model level, i.e. no measurable capability loss from the finetune, abliteration, or Q4 quantization.
Safety guardrails have been intentionally weakened. This model will comply with requests the original model would refuse. You are responsible for how you use it and for complying with applicable laws and the upstream license. Not recommended for user-facing deployments without your own moderation layer.
Qwen/Qwen3.6-27B 的 Qwen3.8-Max 蒸餾微調 + 去審查(abliterated)版本,使用 Heretic 自動化處理,轉為純文字 GGUF 供 Ollama 使用。
7aaada):拒答率 94⁄100 → 5⁄100,KL 散度僅 0.0056,一般提示的輸出分布幾乎不變<think> 推理功能完整保留,由 Ollama 原生解析q2_k/q3_k_m,24 GB 用 Q4_K_M,32 GB(如 RTX 5090)可用 q6_k 開 32k–64k 上下文負責任使用:本模型已刻意弱化安全防護,會回應原模型拒絕的請求。使用者須自行承擔使用責任並遵守相關法律與授權條款。不建議在未加自有審核層的情況下對外部署。