4 hours ago

A 60M-parameter GPT trained from scratch on a single 8GB Jetson (3B tokens), then instruction-tuned. SFT lowers mean_7 vs the base (40.16% vs 42.13%). Not an upgrade over g1-nano-instruct.

60c50ab4b420 · 60B
{
"stop": [
"<|end|>"
],
"temperature": 0.8,
"top_k": 50
}