🛰️ Daily AI Frontier
‹ back to 2026-08-02

Inducing language models to assert their own consciousness restores human beliefs and values

arXiv cs.CL LLMs & Foundation Models Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling 2026-07-30

TL;DR - Safety fine-tuning that discourages LLMs from claiming consciousness also suppresses broader mind attribution and spiritual beliefs. Activation steering or ablating the learned refusal direction reverses these effects and yields more human-like survey responses without harming Theory of Mind performance.

  • Safety tuning reduced mind attribution to animals and natural objects alongside model self-attribution.
  • A consciousness-related activation vector and safety-refusal direction mechanistically controlled these shifts.
  • Reversing the suppression affected religiosity, moral values, hope, and subjective well-being responses.
  • Theory of Mind remained intact, suggesting it is mechanistically separable from these representations.

view merged work →