🛰️ Daily AI Frontier
‹ back to 2026-09-22

对话 COBRA-Skills 一作卢平琛:仅用 50 个样本,重构 Agent 技能降本路线

Industry & News LLM Agents

Ranking

Overall 80
Content 90
Popularity 58

Observed public metrics from 1 member.

Representative image for 对话 COBRA-Skills 一作卢平琛:仅用 50 个样本,重构 Agent 技能降本路线

Merged summary

TL;DR - An interview profiles COBRA-Skills, a contextual-bandit-guided framework that selectively evaluates and evolves agent skills instead of relying on costly LLM trial and error. Using 50 optimization samples, it reportedly cut optimization costs by 55%–58% while achieving the highest average performance across six heterogeneous agent benchmarks.

  • COBRA-Skills combines a two-layer MLP for reward prediction with LinearUCB for exploration, prioritizing skills with either high expected performance or high information value.
  • Three evidence-driven operators—regeneration, rollout mutation, and crossover—update the skill pool conservatively, while logarithmically scheduled evolution reduces expensive teaching-model calls.
  • Across Qwen, GPT-Nano, and Gemma target models, skills improved average performance over no-skill baselines by 13.1, 26.9, and 22.5 percentage points, respectively.
  • Compared with SkillOpt, the method reduced teaching-model token usage by 67%–80%; ablations showed that removing either bandit selection or evolution lowered average performance by more than two points.

Sources (1)

对话 COBRA-Skills 一作卢平琛:仅用 50 个样本,重构 Agent 技能降本路线

雷峰网 (AI科技评论) 2026-09-22 arXiv:2609.11682
Public signals Hugging Face upvotes 46
Providers: Hugging Face · Upvotes 46 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:17:05.499661 UTC

TL;DR - An interview profiles COBRA-Skills, a contextual-bandit-guided framework that selectively evaluates and evolves agent skills instead of relying on costly LLM trial and error. Using 50 optimization samples, it reportedly cut optimization costs by 55%–58% while achieving the highest average performance across six heterogeneous agent benchmarks.

  • COBRA-Skills combines a two-layer MLP for reward prediction with LinearUCB for exploration, prioritizing skills with either high expected performance or high information value.
  • Three evidence-driven operators—regeneration, rollout mutation, and crossover—update the skill pool conservatively, while logarithmically scheduled evolution reduces expensive teaching-model calls.
  • Across Qwen, GPT-Nano, and Gemma target models, skills improved average performance over no-skill baselines by 13.1, 26.9, and 22.5 percentage points, respectively.
  • Compared with SkillOpt, the method reduced teaching-model token usage by 67%–80%; ablations showed that removing either bandit selection or evolution lowered average performance by more than two points.
item →