🛰️ Daily AI Frontier
‹ back to 2026-07-22

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

Research LLMs & Foundation Models

Merged summary

TL;DR - OPIUM is a training-free method that sanitizes activation-steering vectors to reduce safety degradation and excessive refusal. It improves the safety–utility tradeoff by directly optimizing model activations against desired and protected reference behaviors.

  • Uses representation matching across separate prompt sets to preserve intended steering effects while correcting failures.
  • Addresses both utility vectors that weaken safety and refusal vectors that reject benign prompts.
  • Optimizes a replacement steering vector without retraining the underlying model.
  • Outperforms vanilla steering and directional ablation across the evaluated settings.

Sources (1)

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

arXiv cs.LG Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru 2026-07-22 arXiv:2607.19806

TL;DR - OPIUM is a training-free method that sanitizes activation-steering vectors to reduce safety degradation and excessive refusal. It improves the safety–utility tradeoff by directly optimizing model activations against desired and protected reference behaviors.

  • Uses representation matching across separate prompt sets to preserve intended steering effects while correcting failures.
  • Addresses both utility vectors that weaken safety and refusal vectors that reject benign prompts.
  • Optimizes a replacement steering vector without retraining the underlying model.
  • Outperforms vanilla steering and directional ablation across the evaluated settings.
item →