30 Downloads Updated 1 week ago
ollama run siddhu19/swift-15-mtp:i3q_s
Updated 1 week ago
1 week ago
3c59c7efa641 · 12GB
This implementation ports Swift-1.5 Qwen3.8 27B with GSQ-RCO quantization into Ollama-compatible GGUF format, featuring the Mixed-Token-Precision (MTP) inference head for enhanced efficiency during decoding. The architecture retains complete fidelity to the original model while enabling faster generation through mixed-token precision strategies.
Key Architecture Features:
| Tier | Standard GGUF | With MTP Head | Overhead | Precision Profile | Storage Requirements |
|---|---|---|---|---|---|
| IQ3_XXS | 10.09 GB | 10.44 GB | +0.35 GB (~3.5%) | Enhanced precision with mixed-token computation | Moderate VRAM required |
| IQ3_S | 11.77 GB | 12.12 GB | +0.35 GB (~3.5%) | Maximum fidelity with selective precision head | High-end consumer hardware |
The MTP variants add exactly 0.35 GB to both IQ3_XXS and IQ3_S compared to their standard counterparts. This overhead consists of:
The model demonstrates significant efficiency gains over the base Qwen architecture:
The implementation reuses ISTA-DASLab’s GSQ-RCO per-tensor allocation with Swift-specific importance matrix (V1MIX) for quantization. The Q3 variants utilize preserved V1MIX variants of the allocation profiles.