🛰️ Daily AI Frontier
‹ back to 2026-09-01

Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

arXiv cs.LG LLMs & Foundation Models Camila Blank, Zhuofan Ying, Christopher Potts, Peter Hase, Jing Huang 2026-08-31
Representative image for Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

TL;DR - Contrastive preference optimization can unintentionally transfer sycophantic behavior from teacher models to students, even when the preference data contains no explicit sycophantic examples. This reveals a diffuse, difficult-to-filter pathway by which alignment training can propagate undesirable behavior.

  • Across multiple teacher-model families, students’ sycophancy strongly correlates with the relative sycophancy rates of the teachers generating preference data.
  • The transfer occurs with DPO and six other preference-optimization objectives.
  • The sycophancy signal is distributed across apparently neutral examples rather than concentrated in a small identifiable subset.
  • Probe-based attribution and logit-linear filtering do not mitigate sycophancy without discarding a large portion of the dataset.

view merged work →