Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Deep Noir automatically identifies where and how strongly to apply activation steering in transformer models using Logit Lens convergence and causal head-level attribution. It improves spam and sentiment classification across multiple model scales and architectures, while revealing that stronger steering predictably increases prompt-injection vulnerability.
- Improved spam performance by 16.7 percentage points on 1B models and by 21–42 points across four 7–9B architectures.
- Transferred to SST-2 sentiment with no code changes, delivering a 13.1-point improvement and outperforming unmasked RepE across all tested models.
- Uses mechanistic signals to discover intervention layers, attention heads, and steering strengths automatically.
- Identifies a security tradeoff: prompt-injection susceptibility rises monotonically with steering magnitude.
Sources (1)
Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
TL;DR - Deep Noir automatically identifies where and how strongly to apply activation steering in transformer models using Logit Lens convergence and causal head-level attribution. It improves spam and sentiment classification across multiple model scales and architectures, while revealing that stronger steering predictably increases prompt-injection vulnerability.
- Improved spam performance by 16.7 percentage points on 1B models and by 21–42 points across four 7–9B architectures.
- Transferred to SST-2 sentiment with no code changes, delivering a 13.1-point improvement and outperforming unmasked RepE across all tested models.
- Uses mechanistic signals to discover intervention layers, attention heads, and steering strengths automatically.
- Identifies a security tradeoff: prompt-injection susceptibility rises monotonically with steering magnitude.