256 1 month ago

From ornith-1.5:35b, with recommend settings from Hugging Face and 128k context

vision 9b 35b
ollama run smtek/ornith-1.5:9b

Models

View all →

Readme

Ornith-1.5

Notes from ornith-ai/Ornith-1.5-35B-A3B-GGUF and the local Ollama build.

  • Chat template: Qwen chat template (modified, see chat_template.jinja). Uses <|im_start|> / <|im_end|>.
  • Reasoning model: assistant turns open with a thinking … response block before the final answer.
  • Tool calling: emits <tool_call> blocks; use qwen3_xml tool-call parser / qwen3 reasoning parser.
  • Vision: multimodal via mmproj-Ornith-1.5-35B-BF16.gguf (clip projector, 446.57M params).

Recommended sampling

  • General tasks: temperature=0.6, top_p=0.95, top_k=20
  • Reproduce reported benchmarks: temperature=1.0
PARAMETER num_ctx 131072          # 128K context
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
PARAMETER temperature 0.6         # HU recommended value
PARAMETER top_k 20
PARAMETER top_p 0.95
PARAMETER min_p 0
PARAMETER presence_penalty 1.5
PARAMETER repeat_penalty 1
PARAMETER num_predict -1

Usage

ollama run smtek/ornith-1.5:35b

Benchmarks (highlights)

  • SWE-bench Verified 79, SWE-bench Pro 59.6, SWE-bench Multilingual 71.4
  • Terminal-Bench 2.1 (Terminus-2) 67.8
  • GPQA Diamond 89.2, HLE (no tools) 25.6
  • MCP-Atlas 70.2, Toolathlon-Verified 48.7
  • ClawEval 72.5 (evaluated at temp 0.6, 256K ctx)

Notes

  • Requires recent runtimes: Transformers ≥ 5.8.1, vLLM ≥ 0.19.1, SGLang ≥ 0.5.9.
  • Ollama KV cache: use OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 for long contexts.