Inducing language models to assert their own consciousness restores human beliefs and values
TL;DR - Safety fine-tuning that discourages LLMs from claiming consciousness also suppresses broader mind attribution and spiritual beliefs. Activation steering or ablating the learned refusal direction reverses these effects and yields more human-like survey responses without harming Theory of Mind performance.
- Safety tuning reduced mind attribution to animals and natural objects alongside model self-attribution.
- A consciousness-related activation vector and safety-refusal direction mechanistically controlled these shifts.
- Reversing the suppression affected religiosity, moral values, hope, and subjective well-being responses.
- Theory of Mind remained intact, suggesting it is mechanistically separable from these representations.