3 3 days ago

A 62M-parameter GPT trained entirely on a single 8GB-RAM NVIDIA Jetson device, without cloud infrastructure or multi-GPU setups. Both pretraining and instruction fine-tuning occurred within this constraint.

dcd8ef80d92d · 48B
{
"stop": [
"<|end|>",
"</s>"
]
}