Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Merged summary
TL;DR - Off-Context GRPO uses privileged guidance to produce rewarding training rollouts on hard reasoning problems, then importance-corrects updates toward the original unguided objective. It improves mathematical reasoning without meaningful additional cost.
- Addresses RLVR’s zero-signal failure when models cannot generate any correct solutions.
- Generates “off-context” rollouts from prompts containing guidance such as solution prefixes.
- Importance correction avoids the objective mismatch and instability of uncorrected guided training.
- Achieves a 3.9% absolute average gain over vanilla GRPO across standard math benchmarks, equivalent to a 13.8% relative improvement.
Sources (1)
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
TL;DR - Off-Context GRPO uses privileged guidance to produce rewarding training rollouts on hard reasoning problems, then importance-corrects updates toward the original unguided objective. It improves mathematical reasoning without meaningful additional cost.
- Addresses RLVR’s zero-signal failure when models cannot generate any correct solutions.
- Generates “off-context” rollouts from prompts containing guidance such as solution prefixes.
- Importance correction avoids the objective mismatch and instability of uncorrected guided training.
- Achieves a 3.9% absolute average gain over vanilla GRPO across standard math benchmarks, equivalent to a 13.8% relative improvement.