565 Downloads Updated an hour ago
ollama run ornith-1.5:9b
Chirp Chirp! 馃惁 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
Ornith-1.5 extends Ornith-1.0 by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For more details on the task, harness, and rollout reward design, please refer to our blog.