SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - SVR is an oracle-free reinforcement-learning framework that teaches language models to self-verify answers and adaptively allocate test-time reasoning. It improves mathematical reasoning while using fewer refinement turns than fixed-budget approaches.
- Generates an answer, correctness verdict, and confidence score at each turn, stopping only when sufficiently confident.
- Uses ground-truth correctness for training rewards but requires no external verifier or oracle at inference.
- Trains Qwen3.5-2B with GRPO using correctness, calibration, and stop-readiness rewards.
- Achieves 0.563 macro-average accuracy across seven math benchmarks with 2.99 inference turns on average.
Sources (1)
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - SVR is an oracle-free reinforcement-learning framework that teaches language models to self-verify answers and adaptively allocate test-time reasoning. It improves mathematical reasoning while using fewer refinement turns than fixed-budget approaches.
- Generates an answer, correctness verdict, and confidence score at each turn, stopping only when sufficiently confident.
- Uses ground-truth correctness for training rewards but requires no external verifier or oracle at inference.
- Trains Qwen3.5-2B with GRPO using correctness, calibration, and stop-readiness rewards.
- Achieves 0.563 macro-average accuracy across seven math benchmarks with 2.99 inference turns on average.