22 Downloads Updated 5 days ago
ollama run adityakarnam/Ornith-1.5-9B-MLX-distil:int4-mlx
Updated 5 days ago
5 days ago
30db42567aaf · 7.9GB
A distilled Ornith-1.5-9B reasoning model for Apple Silicon, packaged for Ollama’s MLX engine. On MMLU and HumanEval it matches the full-precision bf16 parent at under half the size.
ollama run adityakarnam/Ornith-1.5-9B-MLX-distil
Requires Ollama with the MLX engine (macOS on Apple Silicon). Thinking is on by default.
All rows measured on the same oMLX harness under the same protocol: MMLU on a fixed 1,000-question sample, TruthfulQA MC1 on all 817, HumanEval on all 164, thinking on, greedy decoding.
| Build | Size | MMLU | TruthfulQA | HumanEval |
|---|---|---|---|---|
| This model (Ollama int4) | 7.9 GB | 82.6 | 77.7 | 91.5 |
| Distilled, oQ4 (Hugging Face) | 4.9 GB | 83.5 | 79.0 | 90.8 |
| Ornith-1.5-9B, oQ4 (stock) | 4.9 GB | 78.0 | 80.7 | 87.8 |
| Ornith-1.5-9B, oQ8 (stock) | 8.9 GB | 83.1 | 80.7 | 88.4 |
| Ornith-1.5-9B, bf16 (parent) | 17 GB | 82.6 | 80.5 | 91.5 |
The weights are the distilled model (v1): two-teacher sequence-level distillation from Ornith-1.5-35B-A3B (code) and Qwen3.6-35B-A3B (knowledge and truthfulness), with rejection sampling so only verified-correct teacher traces were trained on. Training was LoRA rank 32 over the bf16 base on one 64 GB M4 Pro, fused afterwards.
Quantization is Ollama’s own MLX recipe applied to those bf16 weights: 152 tensors at int4, 49 at int8 (including the output head and some attention and MLP projections), and 226 kept in bf16 (embeddings, norms, gated-delta parameters), all at group size 64.
This is not the oQ4 build published on Hugging Face. oQ4 promotes 120 modules to 5- and 6-bit, which Ollama’s MLX format cannot represent (it supports int4, int8, nvfp4, mxfp4 and mxfp8), so the oQ4 files will not load here. This build keeps more precision in more places, which is why it is larger and why it ties the parent on two of three benchmarks.
This model’s main weakness, like most quantized reasoning models, is that it sometimes keeps reasoning past the point where more reasoning helps, and runs out of budget before answering.
num_predict low on hard problems. With a 2,048-token cap, the oQ4 build of these weights ran out of budget before answering on 24 of 164 HumanEval problems.</think> token that grows the longer reasoning runs. On the oQ4 build it lifted HumanEval from 130 to 147 of 164 at a 2,048-token budget with no regressions, and it helped every model tested, full precision included. It needs a logits processor, which Ollama does not expose. To use it, run the MLX build with the project’s harness: odistil eval --think-bias 1024:0.02.| Architecture | qwen3_5 hybrid: gated DeltaNet linear attention plus full attention |
| Parameters | 9.4B |
| Context | 262,144 tokens |
| Capabilities | completion, tools, thinking |
| System prompt | You are Ornith, a careful and concise assistant. |
Built by Aditya Karnam Gururaj Rao.