TokenRhythm/ neohorse-1:4b-mlx-8bit

377 3 weeks ago

tools thinking
ollama run TokenRhythm/neohorse-1:4b-mlx-8bit

Details

3 weeks ago

13e8fae32250 · 4.5GB

{ "architectures": [ "Qwen3_5ForCausalLM" ], "attention_bias": false, "attention_dropout": 0.0, "att
{ "eos_token_id": [ 248044, 248046 ] }
{ "version": "1.0", "truncation": null, "padding": null, "added_tokens": [ { "id": 248044, "content"
{ "add_prefix_space": false, "audio_bos_token": "<|audio_start|>", "audio_eos_token": "<|audio_end|>
Apache License Version 2.0, January 2004 http://www.apache.org/licenses/ TERMS AND CONDITIONS FOR US
{ "num_ctx": 4096 }
426 tensors

Readme

NeoHorse-1

Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

GitHub · Technical Report · TokenRhythm · Hugging Face · X

NeoHorse-1 is a family of 4B and 9B causal language models, post-trained by TokenRhythm from Qwen3.5 for text-based agent harnesses, tool use, coding, reasoning, and instruction following. It is an initial prototype on the path toward recursive self-improvement (RSI).

Highlights

  • Path toward RSI: a routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation-selection-update loop.
  • Agentic post-training: routing-guided curriculum SFT and routing-guided on-policy distillation turn execution trajectories into training signal while preserving the execution context of each response.
  • Data quality: exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
  • Original-checkpoint results: the 4B model reports a ten-benchmark macro average of 64.87 versus 58.94 for Qwen3.5-4B; the 9B model reports 69.04 versus 65.60 for Qwen3.5-9B.

Benchmarks

4B results

NeoHorse-1-4B original-checkpoint evaluation

Benchmark Qwen3.5-4B Gemma-4-E4B-it Nanbeige-4.2-3B Agents-A1-4B Spark-X2.5-4B NeoHorse-1-4B Δ vs Qwen3.5-4B
🤖 Agentic
QwenClawBench 38.47 22.98 40.66 43.16 43.52 44.68 +6.21
WorkBuddy Bench 24.62 11.65 21.03 33.37 26.47 34.41 +9.79
PinchBench 71.19 47.60 66.78 75.07 62.37 77.33 +6.14
VitaBench 21.50 5.00 31.50 39.25 37.00 32.00 +10.50
BFCL v4 61.02 47.18 67.28 46.60 63.71 61.79 +0.77
tau2-Bench 84.29 43.60 85.08 81.00 77.72 88.46 +4.17
💻 Coding
HumanEval 87.20 84.76 98.78 92.68 92.07 96.95 +9.75
LiveCodeBench v6 53.71 52.00 72.50* 56.57 54.86 59.43 +5.72
📚 Instruction Following
IFBench 60.33 40.00 55.00 63.33 73.33 65.33 +5.00
IFEval 87.06 74.68 84.47 83.55 91.13 88.35 +1.29
📊 Overall
Ten-benchmark average 58.94 42.95 62.31 61.46 62.22 64.87 +5.93

* Nanbeige-4.2-3B LiveCodeBench v6 result is reported in the corresponding model’s official blog post or technical report.

9B results

NeoHorse-1-9B original-checkpoint evaluation

Benchmark Granite-4.2-8B Qwen3.5-9B Ornith-1.5-9B Gemma-4-12B-it Muse-Glimmer-30B NeoHorse-1-9B Δ vs Qwen3.5-9B
🤖 Agentic
QwenClawBench 37.01 44.04 47.27 43.53 46.11 48.73 +4.69
WorkBuddy Bench 35.07 39.60 29.29 29.65 45.85 40.15 +0.55
PinchBench 56.93 74.55 68.22 58.89 71.35 82.25 +7.70
VitaBench 23.00 31.25 26.75 36.50 48.50 42.25 +11.00
BFCL v4 52.06 64.88 65.03 62.06 53.74 67.43 +2.55
tau2-Bench 62.28 88.04 83.68 59.37 76.64 90.82 +2.78
💻 Coding
HumanEval 96.34 92.68 93.90 100.00 98.17 98.17 +5.49
LiveCodeBench v6 72.00 65.14 47.43 73.14 65.71 65.14 +0.00
📚 Instruction Following
IFBench 78.00 66.33 40.00 77.67 78.67 66.33 +0.00
IFEval 92.98 89.46 71.35 94.27 93.90 89.09 -0.37
📊 Overall
Ten-benchmark average 60.57 65.60 57.29 63.51 67.86 69.04 +3.44

Reported source protocol: SGLang v0.5.17, temperature 1.0, top-p 0.95, top-k 20, min-p 0.0, presence penalty 1.5, repetition penalty 1.0, thinking enabled with enable_thinking=true and force_nonempty_content=true. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; other benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.

License

NeoHorse-1 is released under Apache-2.0. The upstream bases are Qwen3.5-4B and Qwen3.5-9B.

Citation

@misc{neohorse2026,
  title        = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
  author       = {NeoHorse Team},
  year         = {2026},
  howpublished = {arXiv preprint}
}

For questions or issues, use the NeoHorse project repository.