ollama run treyleo16/gpt-2
OpenAI’s GPT-2 Large packaged for Ollama as a Q4_K_M GGUF, with a Q&A prompt template so the base model can answer short questions in chat.
ollama run treyleo16/gpt-2
| Base model | GPT-2 Large (OpenAI, 2019) |
| Architecture | gpt2 |
| Parameters | 774M (Ollama reports 838M) |
| Quantization | Q4_K_M |
| Download size | ~550 MB |
| Context length | 1024 tokens |
| Embedding size | 1280 |
| Capabilities | Text completion |
| License | MIT (inherited from GPT-2) |
GPT-2 Large is a base model. It was never instruction-tuned or trained on chat data, so it predicts the next text and does not follow instructions. The template here frames each turn as a Q: / A: exchange, and stop sequences cut it off after one answer line, so it gives short answers.
Good for: - Short factual questions and one-line answers - Raw text continuation (story starts, article openings) - Experiments, benchmarks, and learning how prompt templates steer a base model - Running on low-end hardware (~20 tok/s on CPU)
Not good for: - Code: its output is not valid, working code - Multi-step instructions, math, or reasoning - Factual accuracy: it writes smoothly and confidently makes things up - Long answers: output stops at the first newline
| Prompt | Response |
|---|---|
| Hi, who are you? | I’m a friendly person that helps people. |
| What is the capital of France? | Paris. |
| What is 2+2? | Two and two are equal to four. |
| What color is the sky? | It’s blue! |
To skip the Q&A template and use GPT-2 as a plain text continuer:
curl http://localhost:11434/api/generate -d '{
"model": "treyleo16/gpt-2",
"raw": true,
"prompt": "Once upon a time in a small mountain town,",
"options": { "stop": [], "num_predict": 200, "temperature": 0.8 }
}'
num_ctx at 1024 or lower. GPT-2 can’t handle longer context.\n stop and keep "Q:". Output gets longer but wanders more.