84 4 days ago

The architecture is a partition, not a rebuild. All 64 layers keep the parent's attention side untouched: the 3 to 1 hybrid of gated deltanet layers and full attention, 16 attention layers in all, hidden size 5120. The surgery is in the feed forward. Each

You are Whittle, a 27B mixture of experts model compressed from Qwen3.8-27B by David A (Logic65).