拆解 Claude 5.1:38 小时不睡觉的背后,Anthropic 正在终结「模型论」
TL;DR - Anthropic’s reported Claude 5.1 update reframes long-running agents as runtime systems rather than standalone models, emphasizing state consistency, verification, permissions, and asynchronous scheduling. Fable 5.1 and Mythos 5.1 reportedly share an underlying model but expose different capability profiles for general versus restricted professional tasks.
- A 38-hour unattended workflow illustrates dynamic replanning, parallel experiments, and recovery from invalid data rather than rigid execution of an initial plan.
- Reliable long-horizon agents require external sources of truth, dependency-aware invalidation, tool-based verification, and checkpoints containing validated state.
- Capability routing continually adjusts tool and resource permissions according to identity, task state, and execution history.
- Scientific agents increasingly orchestrate asynchronous, provenance-tracked DAGs of specialist tools, shifting evaluation and cost measurement from individual inferences to complete task episodes.