31 Downloads Updated 1 week ago
ollama run siddhu19/swift-15:i3q_s
This implementation ports Switf-1.5 Qwen3.8 27B with GSQ-RCO quantization into Ollama-compatible GGUF format. The architecture retains complete fidelity to the original model while enabling efficient inference through mixed-precision quantization.
Key Architecture Features:
| Tier | File Size | Precision Profile | Storage Requirements |
|---|---|---|---|
| IQ3_XXS | 10.09 GB | Enhanced precision for sensitive tasks | Moderate VRAM required |
| IQ3_S | 11.77 GB | Maximum fidelity within constrained sizes | High-end consumer hardware |
The model demonstrates significant efficiency gains over the base Qwen architecture:
The implementation reuses ISTA-DASLab’s GSQ-RCO per-tensor allocation with Swift-specific importance matrix (V1MIX) for quantization. The Q3 variants utilize preserved V1MIX variants of the allocation profiles.