🛰️ Daily AI Frontier
‹ back to 2026-08-17

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

arXiv cs.AI LLMs & Foundation Models Yiqi Liu, Yang Wang, Songxin Wang, Chenghao Xiao, Chenghua Lin 2026-08-16

TL;DR - This paper finds that LLM refusals suppress rather than erase correct answers, which remain recoverable from hidden states. Releasing an answer requires a highly local intervention, while restoring a coherent refusal requires broader changes—an asymmetry important for safety auditing and model steering.

  • Bidirectional activation patching compares matched answering and refusal trajectories in a controlled withholding setup.
  • A single-position patch can release a withheld answer, but reinstating suppression requires patches across multiple positions.
  • Linear probes recover correct answers even during clean refusals, so recoverability does not necessarily imply behavioral control.
  • Average answer-to-refusal displacement vectors distinguish the states geometrically but do not provide a reliable reversible control switch.

view merged work →