🛰️ Daily AI Frontier
‹ back to 2026-09-19

Can Data Attribution Filter Out Subliminal Learning? Not Reliably

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper tests whether gradient-based data attribution can identify and filter training data that transmits hidden behavioral traits through subliminal learning. EK-FAC partially mitigates the effect, but no evaluated method works reliably across models and settings.

  • EK-FAC is the strongest attribution method for token-level filtering, while GradCos and contrastive GradCos provide little benefit.
  • All attribution methods generally underperform the counterfactual-teacher-based divergence-token baseline at token-level filtering.
  • Filtering entire samples is less effective overall, although EK-FAC often provides a stronger signal than divergence tokens in that setting.
  • Performance varies substantially across model–preference combinations, with no consistent explanation for the differences.

Sources (1)

Can Data Attribution Filter Out Subliminal Learning? Not Reliably

arXiv cs.AI Moritz Weckbecker, Sweta Jena, Jonas Müller, Ponnurangam Kumaraguru, Sebastian Lapuschkin, Wojciech Samek, Louis Jaburi, Gonçalo Paulo 2026-09-17 arXiv:2609.20027
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-19 14:12:21.211850 UTC

TL;DR - This paper tests whether gradient-based data attribution can identify and filter training data that transmits hidden behavioral traits through subliminal learning. EK-FAC partially mitigates the effect, but no evaluated method works reliably across models and settings.

  • EK-FAC is the strongest attribution method for token-level filtering, while GradCos and contrastive GradCos provide little benefit.
  • All attribution methods generally underperform the counterfactual-teacher-based divergence-token baseline at token-level filtering.
  • Filtering entire samples is less effective overall, although EK-FAC often provides a stronger signal than divergence tokens in that setting.
  • Performance varies substantially across model–preference combinations, with no consistent explanation for the differences.
item →