王云鹤创业后交出首个模型
TL;DR - TokenRhythm unveiled NeoHorse-1, its first Agent-Native model family, trained on execution traces from a multi-model routing system. The approach turns tool use, failures, recovery paths, and environment feedback into post-training data for more capable and cost-efficient agents.
- NeoHorse-1 comes in 4B and 9B variants and targets tool calling, environment feedback interpretation, error recovery, path adjustment, and task completion.
- Training combines curated public data with Routing Harness traces, using routing-guided curriculum learning and on-policy distillation.
- Agentic post-training raised the 4B model’s reported macro-average score from 58.94 to 64.87, matching or slightly exceeding the 9B base model overall across 10 evaluations.
- The release demonstrates a single execution-to-training improvement cycle, but the company says sustained multi-generation recursive self-improvement remains unproven.