59 1 week ago

NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward recursive self-improvement (RSI).

9b
ollama run maternion/neohorse-1:9b

Details

1 week ago

c59cc7553340 · 5.6GB

qwen35
·
8.95B
·
Q4_K_M
{ "min_p": 0, "presence_penalty": 1.5, "repeat_penalty": 1, "temperature": 0.7,

Readme

NeoHorse-1

Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

Technical Report

NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-9B for text-based agent harnesses, tool use, coding, and instruction following.

These files contain text-only model weights, fine-tuned by TokenRhythm from Qwen3.5-9B.

NeoHorse-1-9B evaluation results

Highlights

  • Path toward RSI: the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation–selection–update loop; extending this loop across successive iterations is the next step toward RSI.
  • Agentic post-training framework: the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
  • Data quality: exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
  • Broad gains: 69.04 macro average across ten benchmarks versus 65.60 for Qwen3.5-9B (+3.44).

Model Details

Property Value
Model family NeoHorse Agent-Native Causal Language Model
Parameters Approximately 9B
Base model Qwen3.5-9B
Post-training Routing-guided agentic post-training
Interface Text input and text output
Context length 262,144 natively and extensible up to 1,010,000 tokens.
Weight format / precision 16-bit (BF16), 8-bit, 5-bit, 4-bit

Evaluation

The results below are from the original checkpoint. The 9B track compares NeoHorse-1-9B with five representative open-weight baselines: Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and Muse-Glimmer-30B. Results cover ten benchmarks and are grouped by capability. Higher is better; Δ is NeoHorse-1-9B minus Qwen3.5-9B. Bold and underline mark the best and second-best results in each benchmark row, respectively; ties share the same formatting.

Benchmark Granite-4.2-8B Qwen3.5-9B Ornith-1.5-9B Gemma-4-12B-it Muse-Glimmer-30B NeoHorse-1-9B Δ vs Qwen3.5-9B
🤖 Agentic
QwenClawBench
37.01
44.04
47.27
43.53
46.11
48.73
+4.69
WorkBuddy Bench
35.07
39.60
29.29
29.65
45.85
40.15
+0.55
PinchBench
56.93
74.55
68.22
58.89
71.35
82.25
+7.70
VitaBench
23.00
31.25
26.75
36.50
48.50
42.25
+11.00
BFCL v4
52.06
64.88
65.03
62.06
53.74
67.43
+2.55
tau2-Bench
62.28
88.04
83.68
59.37
76.64
90.82
+2.78
💻 Coding
HumanEval
96.34
92.68
93.90
100.00
98.17
98.17
+5.49
LiveCodeBench v6
72.00
65.14
47.43
73.14
65.71
65.14
+0.00
📚 Instruction Following
IFBench
78.00
66.33
40.00
77.67
78.67
66.33
+0.00
IFEval
92.98
89.46
71.35
94.27
93.90
89.09
-0.37
📊 Overall
Ten-benchmark average
60.57
65.60
57.29
63.51
67.86
69.04
+3.44

Reported protocol: SGLang v0.5.17 · temperature=1.0 · top_p=0.95 · top_k=20 · min_p=0.0 · presence_penalty=1.5 · repetition_penalty=1.0 · thinking mode enabled with enable_thinking=true and force_nonempty_content=true. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.

Deployment

This repository provides GGUF versions of NeoHorse-1-9B for local use with Ollama. It includes 16-bit (BF16) weights and smaller 8-bit, 5-bit, and 4-bit quantized versions. Quantized versions take up less disk space and use less memory, making the model easier to run on your own hardware.

These are text-only GGUF files with an embedded chat template and no MTP draft head. Use a recent runtime with Qwen3.5 support.

ollama run maternion/neohorse-1:9b
Tag Precision Size
9b 4-bit (Q4_K_M alias) 5.63 GB
9b-q4_K_M 4-bit 5.63 GB
9b-q6_K 6-bit 7.36 GB
9b-q8_0 8-bit 9.53 GB
9b-f16 16-bit (BF16) 17.92 GB

These standard llama.cpp quantizations were generated directly from the BF16 GGUF, without an importance matrix.

License

NeoHorse-1-9B is released under the Apache License 2.0.

The upstream model is Qwen/Qwen3.5-9B. Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm has modified the model through fine-tuning and repackaging for text-only inference.

Citation

@misc{neohorse2026,
  title        = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
  author       = {NeoHorse Team},
  year         = 2026,
  howpublished = {arXiv preprint},
  eprint       = {2609.08183},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url          = {https://arxiv.org/abs/2609.08183}
}

For questions or issue reports, use the NeoHorse project repository.