7 3 days ago

Fast, private code autocomplete: Qwen2.5-Coder-1.5B fine-tuned for fill-in-the-middle with project context. Python, JS, TS.

ollama run thnam02/Kestrel:v0.1

Details

3 days ago

ce0a491c4fd5 · 1.6GB

qwen2
·
1.54B
·
Q8_0
{ "num_ctx": 4096, "temperature": 0.2 }

Readme

Kestrel

A small, fast code autocomplete model that runs entirely on your own computer. Give it the code before and after your cursor, and it fills in the middle.

Kestrel is Qwen2.5-Coder-1.5B fine-tuned for fill-in-the-middle (FIM) completion with project context: snippets from related files in the same project. It was trained on permissively licensed Python, JavaScript and TypeScript code.

  • Size: 1.6 GB, runs on an ordinary laptop
  • Languages: Python, JavaScript, TypeScript
  • Private: no cloud, your code stays on your machine

Kestrel is not a chat model. ollama run thnam02/Kestrel "write a function..." won’t give useful answers. Use it through an editor extension or the raw FIM prompt below.

Quick start

ollama pull thnam02/Kestrel
curl http://localhost:11434/api/generate -d '{
  "model": "thnam02/Kestrel",
  "raw": true,
  "stream": false,
  "prompt": "<|fim_prefix|>def fibonacci(n):\n    <|fim_suffix|>\n\nprint(fibonacci(10))<|fim_middle|>",
  "options": { "num_predict": 64, "temperature": 0.2 }
}'

Response:

if n == 0:
        return 0
    elif n == 1:
        return 1
    else:
        return fibonacci(n - 1) + fibonacci(n - 2)

Always set "raw": true, so Ollama sends the prompt exactly as written.

Prompt format

Single file:

<|fim_prefix|>{code before cursor}<|fim_suffix|>{code after cursor}<|fim_middle|>

With project context (the format Kestrel was trained on, and where it does best):

<|repo_name|>{project name}
<|file_sep|>{path/to/related_file.py}
{related snippet}
<|file_sep|>{path/to/current_file.py}
<|fim_prefix|>{code before cursor}<|fim_suffix|>{code after cursor}<|fim_middle|>

Recommended limits, matching training: about 3,000 characters before the cursor, 1,000 after, and up to 2,000 characters of project context (up to 3 snippets).

Stop generation at any of: <|endoftext|>, <|fim_prefix|>, <|fim_suffix|>, <|fim_middle|>, <|fim_pad|>, <|file_sep|>, <|repo_name|>.

Results

Compared with the original Qwen2.5-Coder-1.5B:

Test Result
Real-world code completion Better: +12.3 points
Gain from adding project context Better: +9.3 points
HumanEval infilling (checking nothing got worse) Same (within noise)

Every “better” result is well above what random variation would produce (about 3 points).

Defaults

Parameter Value
temperature 0.2
num_ctx 4096

Licence

Kestrel is a modified version of Qwen/Qwen2.5-Coder-1.5B, released under the Apache License 2.0. It was fine-tuned for fill-in-the-middle autocomplete with project context, on permissively licensed code from ronantakizawa/github-top-code (Python, JavaScript, TypeScript).