CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution
Ranking
Overall
85
Content
95
Popularity
61
Observed public metrics from 1 member.
Merged summary
TL;DR - CoSkill is a multi-agent reinforcement learning framework that jointly trains an LLM reasoning agent and a learnable meta-skill agent over a hierarchical skill library. This co-adaptation improves sample efficiency, task performance, and wall-clock efficiency compared with prior skill-based and RL baselines.
- The reasoning and meta-skill agents operate cooperatively using a shared model backbone.
- The reasoning agent uses a retrieved task skill and selected child step skills, while task outcomes guide the meta-skill agent in refining those steps.
- CoSkill replaces fixed meta-skill workflows with a learned agent, enabling end-to-end skill evolution alongside policy optimization.
- It achieves 98.4% success on ALFWorld and 90.6% on WebShop, improving over prior baselines by 3.5 and 6.2 percentage points, respectively.
Sources (1)
CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution
Public signals
Hugging Face upvotes 1 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - CoSkill is a multi-agent reinforcement learning framework that jointly trains an LLM reasoning agent and a learnable meta-skill agent over a hierarchical skill library. This co-adaptation improves sample efficiency, task performance, and wall-clock efficiency compared with prior skill-based and RL baselines.
- The reasoning and meta-skill agents operate cooperatively using a shared model backbone.
- The reasoning agent uses a retrieved task skill and selected child step skills, while task outcomes guide the meta-skill agent in refining those steps.
- CoSkill replaces fixed meta-skill workflows with a learned agent, enabling end-to-end skill evolution alongside policy optimization.
- It achieves 98.4% success on ALFWorld and 90.6% on WebShop, improving over prior baselines by 3.5 and 6.2 percentage points, respectively.