🛰️ Daily AI Frontier
‹ back to 2026-09-11

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

arXiv cs.AI LLM Agents Pingchen Lu, Xiangyi Wang, Xiang Li, Jie Mao, Zikun Qu, Junfeng Luo, Yao Shu, Bryan Kian Hsiang Low, Zhongxiang Dai 2026-09-10
Representative image for COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

TL;DR - COBRA-Skills optimizes reusable LLM-agent skills by combining contextual-bandit prioritization with execution-feedback-driven skill evolution. It reports leading average performance across multiple benchmarks while cutting optimization costs and data requirements.

  • Treats skill optimization as budgeted sequential optimization over a dynamically evolving candidate pool.
  • Selectively evaluates promising or informative skills rather than exhaustively testing candidates.
  • Across six agent benchmarks and three target models, it achieves the strongest average performance among compared methods.
  • Reduces optimization cost by 55–58% versus SkillOpt while using 50 unique optimization examples per benchmark.

view merged work →