🛰️ Daily AI Frontier
‹ back to 2026-08-13

从 RLVR 到 RLSVR:开放式任务如何获得可核验奖励

WeChat: 深度学习自然语言处理 LLMs & Foundation Models 2026-08-11
Representative image for 从 RLVR 到 RLSVR:开放式任务如何获得可核验奖励

TL;DR - RLSVR turns open-ended LLM tasks into games with self-verifiable outcomes; its SpyRL method uses an information-asymmetry voting game to train models without external judges. Results suggest improvements in summarization, creative writing, and reasoning, though voting remains only a proxy for quality.

  • Five agents complete one task, with one receiving incomplete input; votes identifying that agent produce rule-verifiable rewards.
  • SpyRL alternates between improving task performance and detecting incomplete-information outputs, creating co-evolving pressure.
  • Qwen3-4B/8B experiments outperformed cited self-improvement baselines, with gains supported by GPT-4o evaluations and limited human blind review.
  • Key uncertainties include proxy gaming, five-agent compute costs, and generalization beyond the tested models and tasks.

view merged work →