352 Downloads Updated 4 days ago
ollama run jikepjikep_16HEX/gemma-4-e4b-nightshift-heretic-uncensored-q4
Updated 4 days ago
4 days ago
4e9df2b05140 ยท 5.3GB ยท
Ultralight Edge Reasoning & Coding Agent Tool Calling โข Thinking โข Rust โข Linux / Windows CLI โข Local AI โข GGUF โข #16HEX Matrix
| Parameter | Value |
|---|---|
| Base Architecture | Google DeepMind Gemma 4 E4B |
| Architecture Profile | Dense + PLE |
| Quantization | i1-Q4_K_M |
| Quantization Method | Importance-matrix optimized (imatrix) |
| Model Size | ~5.3 GB |
| Primary Workloads | Reasoning โข Coding โข Tools โข CLI Agents |
| Modalities | Text โข Tools โข Logical Reasoning |
| Context Window | 32,768 tokens |
| Maximum Generation | 8,192 tokens |
| Deployment Target | Edge AI โข Local AI โข CPU / GPU / iGPU |
Gemma 4 E4B Nightshift Heretic Uncensored is a compact local AI model designed for reasoning, software engineering, tool calling and terminal-oriented workflows.
The i1-Q4_K_M build is focused on efficient local inference while maintaining a practical balance between model size, memory requirements, inference speed and output quality.
The ~5.3 GB model footprint makes this variant particularly suitable for constrained local environments, including laptops, iGPU systems and smaller GPU configurations.
The model is configured for structured reasoning workflows using its supported thinking channel:
<|think|>
This enables an explicit reasoning stage for tasks requiring:
Note: a low-temperature configuration can make generation more consistent, but it does not guarantee factual correctness or eliminate hallucinations.
A primary positioning of this build is tool-oriented local AI.
Suitable workflows include:
The exact tool interface depends on the Ollama runtime and the agent/frontend integrating the model.
The model is positioned for compact, local software-engineering workflows, including:
The modelโs suitability for a programming language comes from its learned capabilities and prompting/tool environment; it should not be interpreted as a guarantee of compiler-level correctness.
This model is intended for integration into local coding-agent environments such as:
Example launch pattern:
ollama launch claude --model jikepjikep_16HEX/gemma-4-e4b-nightshift-heretic-uncensored-q4
ollama launch opencode --model jikepjikep_16HEX/gemma-4-e4b-nightshift-heretic-uncensored-q4
Availability and exact command syntax depend on the installed Ollama version and the respective integration.
Recommended configuration:
temperature : 0.20
top_k : 16
min_p : 0.05
repeat_penalty : 1.10
num_ctx : 32768
num_predict : 8192
| Parameter | Value | Purpose |
|---|---|---|
temperature |
0.20 | Conservative token sampling |
top_k |
16 | Restricts candidate token set |
min_p |
0.05 | Removes low-probability candidates |
repeat_penalty |
1.10 | Reduces repetitive generation |
num_ctx |
32768 | 32K-token working context |
num_predict |
8192 | Maximum generated output |
The configuration is designed to favor:
Precision โ Stability โ Information Density โ Controlled Generation
top_k=16 also provides the direct numerical connection to the 16HEX deployment philosophy:
16 = 2โด
The branding is therefore reflected directly in the sampling configuration without claiming that hexadecimal representation itself changes model computation.
Nightshift Heretic Uncensored identifies this as an experimentally modified model variant intended for a less refusal-oriented interaction profile.
The designation should be understood as a model-variant characteristic, not as a guarantee of:
i1-Q4_K_M is the central deployment format of this build.
The target is an efficient balance between:
Model Size
โ
Memory Usage
โ
Inference Efficiency
โ
Local Deployment
โ
Practical Coding / Agent Workloads
The approximately 5.3 GB model size makes this variant particularly attractive for:
Actual RAM / VRAM requirements depend on context length, runtime, KV-cache configuration and CPU/GPU offloading.
When executed locally through Ollama, inference can remain on the userโs machine instead of requiring a cloud model endpoint.
This makes the model suitable for:
Network privacy ultimately depends on the complete runtime stack, frontend, plugins, tools and external APIs connected to the model.
Gemma 4 E4B Nightshift Heretic Uncensored is optimized for compact local workflows involving:
AI Coding โข Reasoning โข Tool Calling โข CLI Agents โข Rust โข Linux โข Windows โข PowerShell โข JSON โข Automation โข Local AI โข Edge AI โข Developer Tools โข Open-Source AI
Model : Gemma 4 E4B
Variant : Nightshift Heretic Uncensored
Quantization : i1-Q4_K_M
Size : ~5.3 GB
Context : 32,768 tokens
Generation : 8,192 tokens
Thinking : โ
Tool Calling : โ
Coding : โ
CLI Workflows: โ
Edge AI : โ
Vision : โ
Format : GGUF / Ollama
Brand : #16HEX Matrix
#16HEX is the deployment philosophy behind this model:
High information density. Controlled inference. Local execution.
0โ9 โ measurable parameters, configuration and system constraints
AโF โ optimization, reasoning, architecture and engineering synergy
The goal is not maximum token generation.
The goal is:
Less noise โ more signal โ stronger structure โ useful output.
Gemma 4 E4B โข Gemma 4 E4B Uncensored โข Gemma 4 E4B Heretic โข Gemma 4 E4B Ollama โข Gemma 4 E4B Q4_K_M โข i1-Q4_K_M โข Abliterix โข Uncensored AI โข Heretic AI โข Local AI โข Edge AI โข Ollama โข GGUF โข AI Coding Agent โข Coding AI โข Reasoning AI โข Thinking Model โข Tool Calling โข Function Calling โข CLI AI โข Terminal AI โข Rust AI โข Rust Coding โข Linux AI โข Windows AI โข PowerShell AI โข Claude Code โข OpenCode โข OpenClaw โข Hermes Agent โข JSON Tool Calling โข Local Coding Assistant โข Offline AI โข Open Source AI โข Developer AI โข Systems Programming โข 16HEX Matrix
gemma4 gemma-4 e4b 27b uncensored heretic abliterix i1-q4_k_m q4_k_m imatrix ollama gguf local-ai edge-ai reasoning thinking tool-calling function-calling coding coding-agent rust linux windows powershell cli terminal-ai claude-code opencode openclaw hermes-agent json automation developer-ai offline-ai open-source-ai 16hex
This model is an experimental local AI system for reasoning, software engineering, tool calling and AI-agent research.
Generated code and technical conclusions should be reviewed and tested by the human operator before use in production or safety-critical environments.
Gemma 4 E4B Nightshift Heretic Uncensored is a compact, local-first reasoning and coding agent built for users who want practical Tool Calling, Thinking, Rust, Linux, Windows CLI, AI coding agents and Edge AI in an efficient i1-Q4_K_M package.
Small footprint. Strong reasoning. Native tools. Local execution.
๐ 16 HEX MATRIX ๐
If the Nightshift Heretic ๐ models are useful to you and you would like to support 16 HEX Matrix / Eastern IT School, you can buy me a coffee with Monero (XMR).
๐ฃ Monero (XMR)
Network: Mainnet
XMR Address:
44DffaT4GKhWDSRxP1FCPfYJbvUXqciDbgMYxnTnNif9Pm5qP4haCmHh8ePEXxQCQRLKNhhnqW8FgDV9UNah7z5CGcBCBQd
Copy the address above and paste it into your Monero wallet.
๐ช New to Monero?
Official Monero resources:
Official Monero Downloads: https://www.getmonero.org/downloads/
Monero Documentation: https://docs.getmonero.org/
XMR exchange services:
FixedFloat: https://ff.io/
ChangeNOW: https://changenow.io/
โก SIMPLE FLOW
Get XMR โ Copy the address โ Send XMR โ Support local AI development.
Your support helps fund the development, testing, hardware and maintenance of the 16 HEX Matrix / Eastern IT School Nightshift Heretic ๐ local AI model series.
Thank you for supporting independent local AI development. ๐ง โก
#MedicalAI #MedicalVision #Radiology #MedicalImaging #EmergencyTriage #ClinicalAI #Qwen35B #Qwen3 #MoE #VisionAI #Ollama #GGUF #LocalAI #MultimodalAI #16HEX