🛰️ Daily AI Frontier
‹ back to 2026-08-17

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

Research LLMs & Foundation Models

Ranking

Overall 80
Content 95
Popularity 44

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper finds that LLM refusals suppress rather than erase correct answers, which remain recoverable from hidden states. Releasing an answer requires a highly local intervention, while restoring a coherent refusal requires broader changes—an asymmetry important for safety auditing and model steering.

  • Bidirectional activation patching compares matched answering and refusal trajectories in a controlled withholding setup.
  • A single-position patch can release a withheld answer, but reinstating suppression requires patches across multiple positions.
  • Linear probes recover correct answers even during clean refusals, so recoverability does not necessarily imply behavioral control.
  • Average answer-to-refusal displacement vectors distinguish the states geometrically but do not provide a reliable reversible control switch.

Sources (1)

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

arXiv cs.AI Yiqi Liu, Yang Wang, Songxin Wang, Chenghao Xiao, Chenghua Lin 2026-08-16 arXiv:2608.15772
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-15 14:32:27.129953 UTC

TL;DR - This paper finds that LLM refusals suppress rather than erase correct answers, which remain recoverable from hidden states. Releasing an answer requires a highly local intervention, while restoring a coherent refusal requires broader changes—an asymmetry important for safety auditing and model steering.

  • Bidirectional activation patching compares matched answering and refusal trajectories in a controlled withholding setup.
  • A single-position patch can release a withheld answer, but reinstating suppression requires patches across multiple positions.
  • Linear probes recover correct answers even during clean refusals, so recoverability does not necessarily imply behavioral control.
  • Average answer-to-refusal displacement vectors distinguish the states geometrically but do not provide a reliable reversible control switch.
item →