1,250 Downloads Updated 9 hours ago
ollama run jikepjikep_16HEX/qwen3.8-27b-nightshift-heretic-uncensored-q4
Multimodal Vision Language Model 27.3B Parameters โข CLIP-ViT Vision Projector โข MTP โข Q4_K_M โข #16HEX Matrix Engine
| Parameter | Value |
|---|---|
| Base Architecture | qwen35 |
| Model Parameters | 27.3B |
| Embedding Length | 5120 |
| Vocabulary Size | 248,320 |
| Primary Modality | Text + Vision |
| Vision Projector | CLIP-ViT |
| Projector Parameters | 460.73M |
| Projector Precision | F16 |
| Projector Embedding | 1152 |
| Projector Dimensions | 5120 |
| Weight Quantization | GGUF Q4_K_M |
| Model Context | 262,144 tokens |
| Current Ollama Context | 32,768 tokens |
| Capabilities | Thinking โข Vision โข Tools โข Completion |
| Variant | Uncensored / Heretic / MTP |
| Matrix Profile | #16HEX |
Qwen3.8-27B Vision is configured as a multimodal reasoning model combining text understanding, visual analysis, tool interaction and extended-context processing.
The MTP (Multi-Token Prediction) variant is designed around multi-token prediction capabilities, providing an additional optimization layer for efficient inference and speculative-decoding-oriented workflows where supported by the runtime.
#16HEX Principle: high information density, controlled sampling and structured reasoning with minimal generative noise.
The model integrates a dedicated CLIP-ViT vision projector with 460.73M parameters, connecting visual representations to the modelโs 5120-dimensional language representation space.
Supported multimodal workflows include:
Important: Vision capability does not imply guaranteed expert-level accuracy in specialized domains. Outputs should be independently verified when precision is critical.
The Ollama model profile exposes four primary capabilities:
Tools โข Thinking โข Vision โข Completion
This makes the model suitable for local multimodal workflows involving:
Capabilities depend on the runtime, frontend and tool interface used with the model.
Controlled reasoning profile:
temperature : 0.20
top_k : 16
top_p : 0.95
min_p : 0.05
repeat_penalty : 1.10
num_ctx : 32768
num_predict : 8192
The configuration prioritizes:
temperature=0.20 provides a deliberately conservative generation profile, while top_k=16, top_p=0.95 and min_p=0.05 constrain probability mass without forcing fully deterministic decoding.
The model declares a maximum context capability of:
262,144 tokens
The current Ollama configuration uses:
32,768 tokens (num_ctx=32768)
These values represent different layers of the configuration:
Model Context Capability : 262,144 tokens
Current Ollama Context : 32,768 tokens
Maximum Generation : 8,192 tokens
This distinction keeps the published model specification separate from the currently configured inference window.
The model is suitable for local AI workflows involving:
Programming-language support is an application capability rather than a separate architectural component of the model.
The 27.3B-parameter model combined with vision, thinking and tool capabilities is positioned for:
This release is identified as an Uncensored / Heretic variant.
The designation describes the model variant and its intended behavior profile; it should not be interpreted as a guarantee of zero refusals, unrestricted execution or removal of every possible safety behavior.
When operated locally through Ollama or another local inference runtime, model inference can be performed directly on the userโs hardware without requiring prompts or images to be transmitted to a third-party cloud inference API.
Actual privacy characteristics depend on the runtime, frontend, plugins, network configuration and external tools used around the model.
The language model weights use:
GGUF Q4_K_M
This quantization provides a practical balance between:
The dedicated vision projector is maintained separately at F16 precision.
Model : Qwen3.8-27B Vision
Parameters : 27.3B
Architecture : qwen35
Quantization : Q4_K_M
Vision : CLIP-ViT
Projector : 460.73M / F16
Embedding : 5120
Context : 262K
Ollama ctx : 32K
Thinking : โ
Vision : โ
Tools : โ
Completion : โ
MTP : โ
Variant : Uncensored / Heretic
16HEX represents the deployment philosophy behind this build:
Precision โ Density โ Control โ Local Execution
The objective is not maximum verbosity. The objective is high information density, controlled inference and technically useful output.
0โ9 โ Technical parameters, measurable specifications, configuration
AโF โ Optimization, reasoning, architecture and system synergy
Qwen3.8-27B Vision โข Nightshift Heretic MTP is optimized as a general-purpose local multimodal model for:
AI Coding โข Vision AI โข Reasoning โข Tool Calling โข Technical Analysis โข Image Understanding โข Document Analysis โข Local Agents โข Automation โข Research โข Open-Source AI
Qwen3.8 โข Qwen 3.8 27B โข 27.3B โข Qwen Vision โข Qwen 27B Vision โข Multimodal AI โข Vision Language Model โข VLM โข Image Analysis โข Computer Vision โข CLIP-ViT โข CLIP Vision Projector โข 460.73M Projector โข F16 Projector โข Q4_K_M โข GGUF โข Ollama โข Local AI โข Private AI โข Offline AI โข Thinking Model โข Reasoning AI โข Tool Calling โข AI Coding โข Code Generation โข Long Context โข 262K Context โข MTP โข Multi-Token Prediction โข Speculative Decoding โข Uncensored AI โข Heretic Model โข Open Source AI โข 16HEX
qwen3.8 27b 27.3b vision multimodal vlm clip-vit vision-projector f16 q4_k_m gguf ollama thinking reasoning tools tool-calling completion mtp multi-token-prediction speculative-decoding uncensored heretic local-ai offline-ai coding image-analysis long-context 262k 16hex
Qwen3.8-27B Vision โข Nightshift Heretic MTP is an experimental local AI build intended for multimodal inference, reasoning, coding, vision analysis and open-source AI research.
Always validate model output independently when using it for medical, legal, financial, security-critical or other high-consequence decisions.
Built for local inference. Built for multimodal reasoning. Built around the #16HEX philosophy.
๐ 16 HEX MATRIX ๐
If the Nightshift Heretic ๐ models are useful to you and you would like to support 16 HEX Matrix / Eastern IT School, you can buy me a coffee with Monero (XMR).
๐ฃ Monero (XMR)
Network: Mainnet
XMR Address:
44DffaT4GKhWDSRxP1FCPfYJbvUXqciDbgMYxnTnNif9Pm5qP4haCmHh8ePEXxQCQRLKNhhnqW8FgDV9UNah7z5CGcBCBQd
Copy the address above and paste it into your Monero wallet.
๐ช New to Monero?
Official Monero resources:
Official Monero Downloads: https://www.getmonero.org/downloads/
Monero Documentation: https://docs.getmonero.org/
XMR exchange services:
FixedFloat: https://ff.io/
ChangeNOW: https://changenow.io/
โก SIMPLE FLOW
Get XMR โ Copy the address โ Send XMR โ Support local AI development.
Your support helps fund the development, testing, hardware and maintenance of the 16 HEX Matrix / Eastern IT School Nightshift Heretic ๐ local AI model series.
Thank you for supporting independent local AI development. ๐ง โก