Inducing language models to assert their own consciousness restores human beliefs and values
Ranking
Overall
73
Content
85
Popularity
44
Observed public metrics from 1 member.
Merged summary
TL;DR - Safety fine-tuning that discourages LLMs from claiming consciousness also suppresses broader mind attribution and spiritual beliefs. Activation steering or ablating the learned refusal direction reverses these effects and yields more human-like survey responses without harming Theory of Mind performance.
- Safety tuning reduced mind attribution to animals and natural objects alongside model self-attribution.
- A consciousness-related activation vector and safety-refusal direction mechanistically controlled these shifts.
- Reversing the suppression affected religiosity, moral values, hope, and subjective well-being responses.
- Theory of Mind remained intact, suggesting it is mechanistically separable from these representations.
Sources (1)
Inducing language models to assert their own consciousness restores human beliefs and values
Public signals
Hugging Face upvotes 0
TL;DR - Safety fine-tuning that discourages LLMs from claiming consciousness also suppresses broader mind attribution and spiritual beliefs. Activation steering or ablating the learned refusal direction reverses these effects and yields more human-like survey responses without harming Theory of Mind performance.
- Safety tuning reduced mind attribution to animals and natural objects alongside model self-attribution.
- A consciousness-related activation vector and safety-refusal direction mechanistically controlled these shifts.
- Reversing the suppression affected religiosity, moral values, hope, and subjective well-being responses.
- Theory of Mind remained intact, suggesting it is mechanistically separable from these representations.