🛰️ Daily AI Frontier
‹ back to 2026-07-21

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

Research Multimodal & Generative

Ranking

Overall 72
Content 85
Popularity 41

Observed public metrics from 1 member.

Merged summary

TL;DR - ST-Veto is a training-free decoding method for diffusion multimodal LLMs that replaces temporally unstable or weakly image-grounded tokens. It improves multimodal reasoning accuracy by up to 9% without additional training or generation cost.

  • Uses second-order Taylor prediction to identify tokens with unstable confidence across diffusion steps.
  • Measures image-attention mass to filter tokens lacking strong visual grounding.
  • Swaps vetoed tokens with safer candidates during iterative unmasking.
  • Consistently outperforms standard decoding policies and prior VLM reasoning methods across multiple models and benchmarks.

Sources (1)

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

arXiv cs.AI Keuntae Kim, Beomseok Lee, Hyunwoo Kim, Yong Suk Choi 2026-07-20 arXiv:2607.17884
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-13 10:16:24.817256 UTC

TL;DR - ST-Veto is a training-free decoding method for diffusion multimodal LLMs that replaces temporally unstable or weakly image-grounded tokens. It improves multimodal reasoning accuracy by up to 9% without additional training or generation cost.

  • Uses second-order Taylor prediction to identify tokens with unstable confidence across diffusion steps.
  • Measures image-attention mass to filter tokens lacking strong visual grounding.
  • Swaps vetoed tokens with safer candidates during iterative unmasking.
  • Consistently outperforms standard decoding policies and prior VLM reasoning methods across multiple models and benchmarks.
item →