🛰️ Daily AI Frontier
‹ back to 2026-07-20

Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 40

Observed public metrics from 1 member.

Merged summary

TL;DR - This mechanistic study finds that masked diffusion language models implement a direction-symmetric induction circuit for in-context learning. Their advantage comes from accessing context on both sides of a masked token, not from stronger left-context induction.

  • Previous- and next-token attention heads encode local context; later induction heads locate matching contexts and copy answers.
  • The circuit works whether the matching source occurs before or after the target.
  • With only left context, diffusion models do not outperform matched autoregressive models on induction.
  • Diffusion models infer an implicit denoising timestep from the global fraction of masked tokens, without explicit timestep embeddings.

Sources (1)

Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models

arXiv cs.CL Andy Catruna, Emilian Radoi 2026-07-17 arXiv:2607.15893
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-10 02:57:32.941269 UTC

TL;DR - This mechanistic study finds that masked diffusion language models implement a direction-symmetric induction circuit for in-context learning. Their advantage comes from accessing context on both sides of a masked token, not from stronger left-context induction.

  • Previous- and next-token attention heads encode local context; later induction heads locate matching contexts and copy answers.
  • The circuit works whether the matching source occurs before or after the target.
  • With only left context, diffusion models do not outperform matched autoregressive models on induction.
  • Diffusion models infer an implicit denoising timestep from the global fraction of masked tokens, without explicit timestep embeddings.
item →