OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
Ranking
Overall
71
Content
85
Popularity
40
Observed public metrics from 1 member.
Merged summary
TL;DR - OPIUM is a training-free method that sanitizes activation-steering vectors to reduce safety degradation and excessive refusal. It improves the safety–utility tradeoff by directly optimizing model activations against desired and protected reference behaviors.
- Uses representation matching across separate prompt sets to preserve intended steering effects while correcting failures.
- Addresses both utility vectors that weaken safety and refusal vectors that reject benign prompts.
- Optimizes a replacement steering vector without retraining the underlying model.
- Outperforms vanilla steering and directional ablation across the evaluated settings.
Sources (1)
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - OPIUM is a training-free method that sanitizes activation-steering vectors to reduce safety degradation and excessive refusal. It improves the safety–utility tradeoff by directly optimizing model activations against desired and protected reference behaviors.
- Uses representation matching across separate prompt sets to preserve intended steering effects while correcting failures.
- Addresses both utility vectors that weaken safety and refusal vectors that reject benign prompts.
- Optimizes a replacement steering vector without retraining the underlying model.
- Outperforms vanilla steering and directional ablation across the evaluated settings.