基元律动发布模型NeoHorse,探索Harness驱动的RSI路径
TL;DR - TokenRhythm introduced NeoHorse-1, a 4B/9B Agent-Native model family post-trained on tool-use and feedback trajectories generated by routing harnesses. It represents a single-cycle engineering test of recursive self-improvement, using agent failures and environmental feedback to guide subsequent model training.
- Training data captures capability prediction, model routing, responses, tool calls, and environment feedback, filtered for completion, evidence consistency, and error recovery.
- The approach combines routing-informed curriculum design, agent-execution supervision, and on-policy distillation tailored to the student model’s own outputs.
- Across 11 agent, tool-use, coding, and instruction-following benchmarks, the report says NeoHorse 4B had the highest unweighted average among listed 4B models and beat the Qwen3.5-9B base model on five benchmarks.
- Gains were strongest on observable, verifiable workflows; larger models retained advantages in complex state tracking, long-horizon debugging, and failure recovery.