4 hours ago

A 60M-parameter GPT trained from scratch on a single 8GB Jetson with 3B tokens (2× G1). Small gains. Base checkpoint: completes text. Not a G1 upgrade. For chat, use g1-nano-instruct.

42fd7b99c006 · 31B
{
"temperature": 0.8,
"top_k": 50
}