What do Reward Models Memorize?
Merged summary
TL;DR - This paper finds that reward models trained on human preferences memorize easy examples and dataset-specific shortcuts rather than reliably learning contextual response quality. These biases may undermine their ability to evaluate unfamiliar response pairs.
- Memorization is disproportionately allocated to easy, high-margin preference pairs.
- Reward models exploit shortcuts such as model identity and user-sampling strategy.
- They overgeneralize heuristics like response length and compliance to unseen comparisons.
- Discriminative preference training produces biased judges that struggle with context-dependent quality.
Sources (1)
What do Reward Models Memorize?
TL;DR - This paper finds that reward models trained on human preferences memorize easy examples and dataset-specific shortcuts rather than reliably learning contextual response quality. These biases may undermine their ability to evaluate unfamiliar response pairs.
- Memorization is disproportionately allocated to easy, high-margin preference pairs.
- Reward models exploit shortcuts such as model identity and user-sampling strategy.
- They overgeneralize heuristics like response length and compliance to unseen comparisons.
- Discriminative preference training produces biased judges that struggle with context-dependent quality.