Models
Docs
Pricing
Sign in
Download
Models
Download
Docs
Pricing
Sign in
AZERDSQ
/
g2-nano-instruct
Updated
4 hours ago
A 60M-parameter GPT trained from scratch on a single 8GB Jetson (3B tokens), then instruction-tuned. SFT lowers mean_7 vs the base (40.16% vs 42.13%). Not an upgrade over g1-nano-instruct.
A 60M-parameter GPT trained from scratch on a single 8GB Jetson (3B tokens), then instruction-tuned. SFT lowers mean_7 vs the base (40.16% vs 42.13%). Not an upgrade over g1-nano-instruct.
Cancel
Name
1 model
Size / Usage
Context
Input
g2-nano-instruct:latest
053b971d4210
• 279MB • 2K context window •
Text input • 4 hours ago
Text input • 4 hours ago
g2-nano-instruct:latest
279MB
2K
Text
053b971d4210
· 4 hours ago