treyleo16/ gpt-2:latest

10 1 week ago

GPT -2 Large, made by OpenAI. 0.8B parameters.

ollama run treyleo16/gpt-2

Details

1 week ago

5a8b73242b1e · 549MB

gpt2
·
838M
·
Q4_K_M
{{ if .System }}{{ .System }} {{ end }}{{ range .Messages }}{{ if eq .Role "user" }}Q: {{ .Content }
{ "num_ctx": 1024, "num_predict": 128, "repeat_penalty": 1.15, "stop": [ "Q:

Readme

treyleo16/gpt-2

OpenAI’s GPT-2 Large packaged for Ollama as a Q4_K_M GGUF, with a Q&A prompt template so the base model can answer short questions in chat.

ollama run treyleo16/gpt-2

Model details

Base model GPT-2 Large (OpenAI, 2019)
Architecture gpt2
Parameters 774M (Ollama reports 838M)
Quantization Q4_K_M
Download size ~550 MB
Context length 1024 tokens
Embedding size 1280
Capabilities Text completion
License MIT (inherited from GPT-2)

What it is (and isn’t)

GPT-2 Large is a base model. It was never instruction-tuned or trained on chat data, so it predicts the next text and does not follow instructions. The template here frames each turn as a Q: / A: exchange, and stop sequences cut it off after one answer line, so it gives short answers.

Good for: - Short factual questions and one-line answers - Raw text continuation (story starts, article openings) - Experiments, benchmarks, and learning how prompt templates steer a base model - Running on low-end hardware (~20 tok/s on CPU)

Not good for: - Code: its output is not valid, working code - Multi-step instructions, math, or reasoning - Factual accuracy: it writes smoothly and confidently makes things up - Long answers: output stops at the first newline

Example output

Prompt Response
Hi, who are you? I’m a friendly person that helps people.
What is the capital of France? Paris.
What is 2+2? Two and two are equal to four.
What color is the sky? It’s blue!

Raw completion mode

To skip the Q&A template and use GPT-2 as a plain text continuer:

curl http://localhost:11434/api/generate -d '{
  "model": "treyleo16/gpt-2",
  "raw": true,
  "prompt": "Once upon a time in a small mountain town,",
  "options": { "stop": [], "num_predict": 200, "temperature": 0.8 }
}'

Tips

  • Keep num_ctx at 1024 or lower. GPT-2 can’t handle longer context.
  • For longer answers, remove the \n stop and keep "Q:". Output gets longer but wanders more.
  • Lower temperature (0.3–0.5) gives more on-topic answers. Higher (0.8–1.0) works better for creative text.

Credits

  • Base model: GPT-2 by OpenAI, “Language Models are Unsupervised Multitask Learners” (Radford et al., 2019)
  • Packaging, template, and tuning: treyleo16