3 3 days ago

A 62M-parameter GPT trained entirely on a single 8GB-RAM NVIDIA Jetson device, without cloud infrastructure or multi-GPU setups. Both pretraining and instruction fine-tuning occurred within this constraint.

2af71558e438 · 10B
Apache-2.0