🛰️ Daily AI Frontier
‹ back to 2026-09-20

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Research LLMs & Foundation Models

Ranking

Overall 78
Content 90
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Merged summary

TL;DR - Deep Noir automatically identifies where and how strongly to apply activation steering in transformer models using Logit Lens convergence and causal head-level attribution. It improves spam and sentiment classification across multiple model scales and architectures, while revealing that stronger steering predictably increases prompt-injection vulnerability.

  • Improved spam performance by 16.7 percentage points on 1B models and by 21–42 points across four 7–9B architectures.
  • Transferred to SST-2 sentiment with no code changes, delivering a 13.1-point improvement and outperforming unmasked RepE across all tested models.
  • Uses mechanistic signals to discover intervention layers, attention heads, and steering strengths automatically.
  • Identifies a security tradeoff: prompt-injection susceptibility rises monotonically with steering magnitude.

Sources (1)

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

arXiv cs.AI Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis 2026-09-17 arXiv:2609.20722
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:15:09.821829 UTC

TL;DR - Deep Noir automatically identifies where and how strongly to apply activation steering in transformer models using Logit Lens convergence and causal head-level attribution. It improves spam and sentiment classification across multiple model scales and architectures, while revealing that stronger steering predictably increases prompt-injection vulnerability.

  • Improved spam performance by 16.7 percentage points on 1B models and by 21–42 points across four 7–9B architectures.
  • Transferred to SST-2 sentiment with no code changes, delivering a 13.1-point improvement and outperforming unmasked RepE across all tested models.
  • Uses mechanistic signals to discover intervention layers, attention heads, and steering strengths automatically.
  • Identifies a security tradeoff: prompt-injection susceptibility rises monotonically with steering magnitude.
item →