30 1 week ago

Compact, mixed-precision GGUF quantizations of Swift 1.5 Qwen3.8-27B, with Swift-specific refinement using the per-tensor allocations from ISTA-DASLab's GSQ-RCO release

ollama run siddhu19/swift-15-mtp:i3q_s

Details

1 week ago

3c59c7efa641 · 12GB

qwen35
·
27.3B
·
BF16
Swift Open License v1.0 TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION 1. Definitions.
{ "draft_num_predict": 3, "min_p": 0, "num_ctx": 262144, "repeat_penalty": 1, "t

Readme

Model Overview

This implementation ports Swift-1.5 Qwen3.8 27B with GSQ-RCO quantization into Ollama-compatible GGUF format, featuring the Mixed-Token-Precision (MTP) inference head for enhanced efficiency during decoding. The architecture retains complete fidelity to the original model while enabling faster generation through mixed-token precision strategies.

Key Architecture Features:

  • Based on Qwen3.8 27B with Swift-specific post-training focused on long-horizon, agentic and coding tasks
  • Uses 58.5% fewer thinking tokens than base model
  • Demonstrates 9.18× speed-up on several coding and reasoning tasks compared to base Qwen
  • MTP head enables selective precision computation during inference for improved throughput

Available Quantization Tiers: MTP Overhead Analysis

Tier Standard GGUF With MTP Head Overhead Precision Profile Storage Requirements
IQ3_XXS 10.09 GB 10.44 GB +0.35 GB (~3.5%) Enhanced precision with mixed-token computation Moderate VRAM required
IQ3_S 11.77 GB 12.12 GB +0.35 GB (~3.5%) Maximum fidelity with selective precision head High-end consumer hardware

MTP Overhead Breakdown:

The MTP variants add exactly 0.35 GB to both IQ3_XXS and IQ3_S compared to their standard counterparts. This overhead consists of:

  • 15 additional head tensors (compared to 851 model tensors in the base GGUF) [source]
  • Each MTP dump contains matching refined model tensors plus the specialized inference head for mixed-token precision decoding
  • The head tensors enable runtime support for selective computation patterns without altering the underlying weight quantization

Performance Characteristics from Swift-1.5

The model demonstrates significant efficiency gains over the base Qwen architecture:

  • 0.35% accuracy improvement over base model while using significantly fewer tokens
  • KLD (Kullback-Leibler divergence) measurements confirm improved fidelity compared to standard quantizations across multiple domains including prose, code, math text, and multilingual capabilities

Quantization Details from Source

The implementation reuses ISTA-DASLab’s GSQ-RCO per-tensor allocation with Swift-specific importance matrix (V1MIX) for quantization. The Q3 variants utilize preserved V1MIX variants of the allocation profiles.