siddhu19/ swift-15:i3q_xxs

31 1 week ago

Compact, mixed-precision GGUF quantizations of Swift 1.5 Qwen3.8-27B, with Swift-specific refinement using the per-tensor allocations from ISTA-DASLab's GSQ-RCO release.

ollama run siddhu19/swift-15:i3q_xxs

Details

1 week ago

5bba78a76780 · 10GB

qwen35
·
26.9B
·
IQ3_XXS
Swift Open License v1.0 TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION 1. Definitions.

Readme

Model Overview

This implementation ports Switf-1.5 Qwen3.8 27B with GSQ-RCO quantization into Ollama-compatible GGUF format. The architecture retains complete fidelity to the original model while enabling efficient inference through mixed-precision quantization.

Key Architecture Features:

  • Based on Qwen3.8 27B with Swift-specific post-training focused on long-horizon, agentic and coding tasks
  • Uses 58.5% fewer thinking tokens than base model
  • Demonstrates 9.18× speed-up on several coding and reasoning tasks compared to base Qwen

Available Quantization Tiers (Q3 Variants Only)

Tier File Size Precision Profile Storage Requirements
IQ3_XXS 10.09 GB Enhanced precision for sensitive tasks Moderate VRAM required
IQ3_S 11.77 GB Maximum fidelity within constrained sizes High-end consumer hardware

Performance Characteristics from Swift-1.5

The model demonstrates significant efficiency gains over the base Qwen architecture:

  • 0.35% accuracy improvement over base model while using significantly fewer tokens
  • KLD (Kullback-Leibler divergence) measurements confirm improved fidelity compared to standard quantizations across multiple domains including prose, code, math text, and multilingual capabilities

Quantization Details from Source

The implementation reuses ISTA-DASLab’s GSQ-RCO per-tensor allocation with Swift-specific importance matrix (V1MIX) for quantization. The Q3 variants utilize preserved V1MIX variants of the allocation profiles.