🛰️ Daily AI Frontier
‹ back to 2026-08-28

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

Research Mechanistic Interpretability

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Representative image for Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

Merged summary

TL;DR - Circuit Condensation post-trains models so a target behavior is carried by a smaller causal circuit, making mechanistic explanations easier to inspect and verify. Across eight models and four behaviors, it reduced circuits by 8.1× on average versus the strongest frozen-model baseline while preserving task performance and general capabilities.

  • Each round prunes low-attribution edges and trains a low-rank adapter to reproduce the original behavior through the remaining graph.
  • Condensed circuits were smaller in 30 of 32 settings, with reductions reaching 316×; control searches without weight updates produced larger circuits in 29 of 32 settings.
  • Exhaustive subset testing found 11 of 19 circuits irreducible, while pair ablations showed that edge effects can depend on one another.
  • For indirect object identification, condensation isolated 24 attention heads, including 17 with documented roles, versus 61 heads in the matched frozen circuit.

Sources (1)

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

arXiv cs.LG Sai Adith Senthil Kumar 2026-08-27 arXiv:2608.27254
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-14 14:15:24.830858 UTC

TL;DR - Circuit Condensation post-trains models so a target behavior is carried by a smaller causal circuit, making mechanistic explanations easier to inspect and verify. Across eight models and four behaviors, it reduced circuits by 8.1× on average versus the strongest frozen-model baseline while preserving task performance and general capabilities.

  • Each round prunes low-attribution edges and trains a low-rank adapter to reproduce the original behavior through the remaining graph.
  • Condensed circuits were smaller in 30 of 32 settings, with reductions reaching 316×; control searches without weight updates produced larger circuits in 29 of 32 settings.
  • Exhaustive subset testing found 11 of 19 circuits irreducible, while pair ablations showed that edge effects can depend on one another.
  • For indirect object identification, condensation isolated 24 attention heads, including 17 with documented roles, versus 61 heads in the matched frozen circuit.
item →