94 5 days ago

The architecture is a partition, not a rebuild. All 64 layers keep the parent's attention side untouched: the 3 to 1 hybrid of gated deltanet layers and full attention, 16 attention layers in all, hidden size 5120. The surgery is in the feed forward. Each

{
"temperature": 0.7,
"top_k": 20,
"top_p": 0.8
}