🛰️ Daily AI Frontier
‹ back to 2026-09-20

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

arXiv cs.AI LLMs & Foundation Models Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis 2026-09-17
Representative image for Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

TL;DR - Deep Noir automatically identifies where and how strongly to apply activation steering in transformer models using Logit Lens convergence and causal head-level attribution. It improves spam and sentiment classification across multiple model scales and architectures, while revealing that stronger steering predictably increases prompt-injection vulnerability.

  • Improved spam performance by 16.7 percentage points on 1B models and by 21–42 points across four 7–9B architectures.
  • Transferred to SST-2 sentiment with no code changes, delivering a 13.1-point improvement and outperforming unmasked RepE across all tested models.
  • Uses mechanistic signals to discover intervention layers, attention heads, and steering strengths automatically.
  • Identifies a security tradeoff: prompt-injection susceptibility rises monotonically with steering magnitude.

view merged work →