Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
Ranking
Overall
80
Content
95
Popularity
44
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper finds that LLM refusals suppress rather than erase correct answers, which remain recoverable from hidden states. Releasing an answer requires a highly local intervention, while restoring a coherent refusal requires broader changes—an asymmetry important for safety auditing and model steering.
- Bidirectional activation patching compares matched answering and refusal trajectories in a controlled withholding setup.
- A single-position patch can release a withheld answer, but reinstating suppression requires patches across multiple positions.
- Linear probes recover correct answers even during clean refusals, so recoverability does not necessarily imply behavioral control.
- Average answer-to-refusal displacement vectors distinguish the states geometrically but do not provide a reliable reversible control switch.
Sources (1)
Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
Public signals
Hugging Face upvotes 0
TL;DR - This paper finds that LLM refusals suppress rather than erase correct answers, which remain recoverable from hidden states. Releasing an answer requires a highly local intervention, while restoring a coherent refusal requires broader changes—an asymmetry important for safety auditing and model steering.
- Bidirectional activation patching compares matched answering and refusal trajectories in a controlled withholding setup.
- A single-position patch can release a withheld answer, but reinstating suppression requires patches across multiple positions.
- Linear probes recover correct answers even during clean refusals, so recoverability does not necessarily imply behavioral control.
- Average answer-to-refusal displacement vectors distinguish the states geometrically but do not provide a reliable reversible control switch.