从 RLVR 到 RLSVR:开放式任务如何获得可核验奖励
Ranking
Overall
83
Content
90
Popularity
68
Observed public metrics from 1 member.
Merged summary
TL;DR - RLSVR turns open-ended LLM tasks into games with self-verifiable outcomes; its SpyRL method uses an information-asymmetry voting game to train models without external judges. Results suggest improvements in summarization, creative writing, and reasoning, though voting remains only a proxy for quality.
- Five agents complete one task, with one receiving incomplete input; votes identifying that agent produce rule-verifiable rewards.
- SpyRL alternates between improving task performance and detecting incomplete-information outputs, creating co-evolving pressure.
- Qwen3-4B/8B experiments outperformed cited self-improvement baselines, with gains supported by GPT-4o evaluations and limited human blind review.
- Key uncertainties include proxy gaming, five-agent compute costs, and generalization beyond the tested models and tasks.
Sources (1)
从 RLVR 到 RLSVR:开放式任务如何获得可核验奖励
Public signals
Hugging Face upvotes 106
TL;DR - RLSVR turns open-ended LLM tasks into games with self-verifiable outcomes; its SpyRL method uses an information-asymmetry voting game to train models without external judges. Results suggest improvements in summarization, creative writing, and reasoning, though voting remains only a proxy for quality.
- Five agents complete one task, with one receiving incomplete input; votes identifying that agent produce rule-verifiable rewards.
- SpyRL alternates between improving task performance and detecting incomplete-information outputs, creating co-evolving pressure.
- Qwen3-4B/8B experiments outperformed cited self-improvement baselines, with gains supported by GPT-4o evaluations and limited human blind review.
- Key uncertainties include proxy gaming, five-agent compute costs, and generalization beyond the tested models and tasks.