🛰️ Daily AI Frontier
‹ back to 2026-09-07

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

Research LLM Agents

Ranking

Overall 85
Content 95
Popularity 61

Observed public metrics from 1 member.

Merged summary

TL;DR - CoSkill is a multi-agent reinforcement learning framework that jointly trains an LLM reasoning agent and a learnable meta-skill agent over a hierarchical skill library. This co-adaptation improves sample efficiency, task performance, and wall-clock efficiency compared with prior skill-based and RL baselines.

  • The reasoning and meta-skill agents operate cooperatively using a shared model backbone.
  • The reasoning agent uses a retrieved task skill and selected child step skills, while task outcomes guide the meta-skill agent in refining those steps.
  • CoSkill replaces fixed meta-skill workflows with a learned agent, enabling end-to-end skill evolution alongside policy optimization.
  • It achieves 98.4% success on ALFWorld and 90.6% on WebShop, improving over prior baselines by 3.5 and 6.2 percentage points, respectively.

Sources (1)

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

arXiv cs.AI Jinyuan Feng, Dongmin Li, Yiqun Chen, Yang Gao, Xing Chen, Huimu Wang, Zhiqiang Pu 2026-09-04 arXiv:2609.04865
Public signals Hugging Face upvotes 1 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:23:29.549343 UTC

TL;DR - CoSkill is a multi-agent reinforcement learning framework that jointly trains an LLM reasoning agent and a learnable meta-skill agent over a hierarchical skill library. This co-adaptation improves sample efficiency, task performance, and wall-clock efficiency compared with prior skill-based and RL baselines.

  • The reasoning and meta-skill agents operate cooperatively using a shared model backbone.
  • The reasoning agent uses a retrieved task skill and selected child step skills, while task outcomes guide the meta-skill agent in refining those steps.
  • CoSkill replaces fixed meta-skill workflows with a learned agent, enabling end-to-end skill evolution alongside policy optimization.
  • It achieves 98.4% success on ALFWorld and 90.6% on WebShop, improving over prior baselines by 3.5 and 6.2 percentage points, respectively.
item →