22 5 days ago

tools thinking
ollama run adityakarnam/Ornith-1.5-9B-MLX-distil

Applications

Claude Code
Claude Code ollama launch claude --model adityakarnam/Ornith-1.5-9B-MLX-distil
OpenCode
OpenCode ollama launch opencode --model adityakarnam/Ornith-1.5-9B-MLX-distil
Hermes Agent
Hermes Agent ollama launch hermes --model adityakarnam/Ornith-1.5-9B-MLX-distil
OpenClaw
OpenClaw ollama launch openclaw --model adityakarnam/Ornith-1.5-9B-MLX-distil

Models

View all →

Readme

Ornith-1.5-9B-MLX-distil

A distilled Ornith-1.5-9B reasoning model for Apple Silicon, packaged for Ollama’s MLX engine. On MMLU and HumanEval it matches the full-precision bf16 parent at under half the size.

ollama run adityakarnam/Ornith-1.5-9B-MLX-distil

Requires Ollama with the MLX engine (macOS on Apple Silicon). Thinking is on by default.

Benchmarks

All rows measured on the same oMLX harness under the same protocol: MMLU on a fixed 1,000-question sample, TruthfulQA MC1 on all 817, HumanEval on all 164, thinking on, greedy decoding.

Build Size MMLU TruthfulQA HumanEval
This model (Ollama int4) 7.9 GB 82.6 77.7 91.5
Distilled, oQ4 (Hugging Face) 4.9 GB 83.5 79.0 90.8
Ornith-1.5-9B, oQ4 (stock) 4.9 GB 78.0 80.7 87.8
Ornith-1.5-9B, oQ8 (stock) 8.9 GB 83.1 80.7 88.4
Ornith-1.5-9B, bf16 (parent) 17 GB 82.6 80.5 91.5
  • Against the stock 4-bit build: +4.6 MMLU and +3.7 HumanEval.
  • Against the bf16 parent: identical on MMLU (826⁄1000) and HumanEval (150⁄164) at 46% of the size.
  • TruthfulQA is 2.8 points below the parent. Every distilled build shows this, and the cause is fewer “I don’t know” answers; see the retrospective.
  • The difference from the oQ4 build is within about one standard error on all three benchmarks.

What this build is

The weights are the distilled model (v1): two-teacher sequence-level distillation from Ornith-1.5-35B-A3B (code) and Qwen3.6-35B-A3B (knowledge and truthfulness), with rejection sampling so only verified-correct teacher traces were trained on. Training was LoRA rank 32 over the bf16 base on one 64 GB M4 Pro, fused afterwards.

Quantization is Ollama’s own MLX recipe applied to those bf16 weights: 152 tensors at int4, 49 at int8 (including the output head and some attention and MLP projections), and 226 kept in bf16 (embeddings, norms, gated-delta parameters), all at group size 64.

This is not the oQ4 build published on Hugging Face. oQ4 promotes 120 modules to 5- and 6-bit, which Ollama’s MLX format cannot represent (it supports int4, int8, nvfp4, mxfp4 and mxfp8), so the oQ4 files will not load here. This build keeps more precision in more places, which is why it is larger and why it ties the parent on two of three benchmarks.

Getting the most out of it

This model’s main weakness, like most quantized reasoning models, is that it sometimes keeps reasoning past the point where more reasoning helps, and runs out of budget before answering.

  • Give it room. Don’t cap num_predict low on hard problems. With a 2,048-token cap, the oQ4 build of these weights ran out of budget before answering on 24 of 164 HumanEval problems.
  • Soft HALT is not available in Ollama. The project’s fix is a small bias on the </think> token that grows the longer reasoning runs. On the oQ4 build it lifted HumanEval from 130 to 147 of 164 at a 2,048-token budget with no regressions, and it helped every model tested, full precision included. It needs a logits processor, which Ollama does not expose. To use it, run the MLX build with the project’s harness: odistil eval --think-bias 1024:0.02.

Parameters

Architecture qwen3_5 hybrid: gated DeltaNet linear attention plus full attention
Parameters 9.4B
Context 262,144 tokens
Capabilities completion, tools, thinking
System prompt You are Ornith, a careful and concise assistant.

Links

Built by Aditya Karnam Gururaj Rao.