🛰️ Daily AI Frontier
‹ back to 2026-07-28

What do Reward Models Memorize?

arXiv cs.LG LLMs & Foundation Models Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova 2026-07-27

TL;DR - This paper finds that reward models trained on human preferences memorize easy examples and dataset-specific shortcuts rather than reliably learning contextual response quality. These biases may undermine their ability to evaluate unfamiliar response pairs.

  • Memorization is disproportionately allocated to easy, high-margin preference pairs.
  • Reward models exploit shortcuts such as model identity and user-sampling strategy.
  • They overgeneralize heuristics like response length and compliance to unseen comparisons.
  • Discriminative preference training produces biased judges that struggle with context-dependent quality.

view merged work →