🛰️ Daily AI Frontier
‹ back to 2026-07-30

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

Research Efficiency & Systems

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Representative image for WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

Merged summary

TL;DR - WIDE dynamically prunes attention-head and FFN-channel groups for each token, improving LLM inference efficiency while retaining more quality than coarse-grained pruning. Its kernel co-design delivers practical prefill and decoding acceleration.

  • Supports token-level dynamic width pruning during both prefill and decoding.
  • Uses a two-stage differentiable training pipeline to learn token-wise sparse execution.
  • At 50% sparsity, reports a 55.1% performance boost over dynamic depth pruning under calibration-only settings.
  • Achieves up to 1.98× prefill and 4.95× decoding kernel speedups, with 1.68× and 1.55× end-to-end acceleration.

Sources (1)

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

arXiv cs.AI Haozhe Hu, Hao Wu, Peiran Yin, Chao Han, Yunpu Ma, Xiaoyu Shen 2026-07-30 arXiv:2607.28418
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-26 14:40:53.636657 UTC

TL;DR - WIDE dynamically prunes attention-head and FFN-channel groups for each token, improving LLM inference efficiency while retaining more quality than coarse-grained pruning. Its kernel co-design delivers practical prefill and decoding acceleration.

  • Supports token-level dynamic width pruning during both prefill and decoding.
  • Uses a two-stage differentiable training pipeline to learn token-wise sparse execution.
  • At 50% sparsity, reports a 55.1% performance boost over dynamic depth pruning under calibration-only settings.
  • Achieves up to 1.98× prefill and 4.95× decoding kernel speedups, with 1.68× and 1.55× end-to-end acceleration.
item →