ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
TL;DR - ContinualSkillBench evaluates whether LLM agents can turn sequential task experience into reusable skills. Agents improve over time, but much of the gain appears to come from contextual adaptation rather than robust skill consolidation.
- The benchmark spans five domains, each with 100 increasingly difficult, interconnected subtasks.
- Sequential execution generally improves performance, with substantial variation across models and domains.
- Explicit skill maintenance performs similarly to in-context learning on average, but helps with reusable procedures and precise outputs.
- Less capable models accumulate larger, more fragmented sets of task-specific skills.