What do Reward Models Memorize?
Ranking
Overall
79
Content
95
Popularity
41
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper finds that reward models trained on human preferences memorize easy examples and dataset-specific shortcuts rather than reliably learning contextual response quality. These biases may undermine their ability to evaluate unfamiliar response pairs.
- Memorization is disproportionately allocated to easy, high-margin preference pairs.
- Reward models exploit shortcuts such as model identity and user-sampling strategy.
- They overgeneralize heuristics like response length and compliance to unseen comparisons.
- Discriminative preference training produces biased judges that struggle with context-dependent quality.
Sources (1)
What do Reward Models Memorize?
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper finds that reward models trained on human preferences memorize easy examples and dataset-specific shortcuts rather than reliably learning contextual response quality. These biases may undermine their ability to evaluate unfamiliar response pairs.
- Memorization is disproportionately allocated to easy, high-margin preference pairs.
- Reward models exploit shortcuts such as model identity and user-sampling strategy.
- They overgeneralize heuristics like response length and compliance to unseen comparisons.
- Discriminative preference training produces biased judges that struggle with context-dependent quality.