🛰️ Daily AI Frontier
‹ back to 2026-09-19

dQwen3.5: Hybrid-Attention Diffusion Language Models

Research LLMs & Foundation Models

Ranking

Overall 84
Content 95
Popularity 59

Observed public metrics from 1 member.

Merged summary

TL;DR - dQwen3.5 adapts Qwen3.5’s hybrid attention–RNN architecture into diffusion language models ranging from 0.8B to 9B parameters. The results suggest hybrid autoregressive backbones can enable efficient diffusion-model adaptation despite their structurally causal RNN layers.

  • The family spans 0.8B, 2B, 4B, and 9B parameter scales.
  • Hybrid models reach a given training loss using roughly half the tokens required by a full-attention control.
  • Across scales, dQwen3.5 exhibits any-order decoding behavior similar to full-attention diffusion language models.
  • The adapted models perform strongly with parallel decoding.

Sources (1)

dQwen3.5: Hybrid-Attention Diffusion Language Models

arXiv cs.CL Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai 2026-09-17 arXiv:2609.20751
Public signals Hugging Face upvotes 1
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:18:00.424486 UTC

TL;DR - dQwen3.5 adapts Qwen3.5’s hybrid attention–RNN architecture into diffusion language models ranging from 0.8B to 9B parameters. The results suggest hybrid autoregressive backbones can enable efficient diffusion-model adaptation despite their structurally causal RNN layers.

  • The family spans 0.8B, 2B, 4B, and 9B parameter scales.
  • Hybrid models reach a given training loss using roughly half the tokens required by a full-attention control.
  • Across scales, dQwen3.5 exhibits any-order decoding behavior similar to full-attention diffusion language models.
  • The adapted models perform strongly with parallel decoding.
item →