OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
TL;DR - OPIUM is a training-free method that sanitizes activation-steering vectors to reduce safety degradation and excessive refusal. It improves the safety–utility tradeoff by directly optimizing model activations against desired and protected reference behaviors.
- Uses representation matching across separate prompt sets to preserve intended steering effects while correcting failures.
- Addresses both utility vectors that weaken safety and refusal vectors that reject benign prompts.
- Optimizes a replacement steering vector without retraining the underlying model.
- Outperforms vanilla steering and directional ablation across the evaluated settings.