Can Data Attribution Filter Out Subliminal Learning? Not Reliably
Ranking
Overall
78
Content
95
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper tests whether gradient-based data attribution can identify and filter training data that transmits hidden behavioral traits through subliminal learning. EK-FAC partially mitigates the effect, but no evaluated method works reliably across models and settings.
- EK-FAC is the strongest attribution method for token-level filtering, while GradCos and contrastive GradCos provide little benefit.
- All attribution methods generally underperform the counterfactual-teacher-based divergence-token baseline at token-level filtering.
- Filtering entire samples is less effective overall, although EK-FAC often provides a stronger signal than divergence tokens in that setting.
- Performance varies substantially across model–preference combinations, with no consistent explanation for the differences.
Sources (1)
Can Data Attribution Filter Out Subliminal Learning? Not Reliably
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper tests whether gradient-based data attribution can identify and filter training data that transmits hidden behavioral traits through subliminal learning. EK-FAC partially mitigates the effect, but no evaluated method works reliably across models and settings.
- EK-FAC is the strongest attribution method for token-level filtering, while GradCos and contrastive GradCos provide little benefit.
- All attribution methods generally underperform the counterfactual-teacher-based divergence-token baseline at token-level filtering.
- Filtering entire samples is less effective overall, although EK-FAC often provides a stronger signal than divergence tokens in that setting.
- Performance varies substantially across model–preference combinations, with no consistent explanation for the differences.