🛰️ Daily AI Frontier
‹ back to 2026-07-22

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Research Multimodal & Generative

Ranking

Overall 66
Content 75
Popularity 44

Observed public metrics from 1 member.

Merged summary

TL;DR - IGGT4D is a streaming Transformer that jointly reconstructs dynamic 4D scenes and maintains object identities from continuous video. It advances online spatial intelligence by unifying camera motion, geometry, and temporally consistent instances over long sequences.

  • Processes frames sequentially using causal spatial-temporal modeling and reused historical context.
  • Incrementally tracks camera motion, scene geometry, and object identity in one representation.
  • Introduces InsScene4D-147K, covering real, synthetic, static, and dynamic scenes with geometry-guided annotations.
  • Outperforms reported streaming baselines across reconstruction, pose estimation, instance tracking, and open-vocabulary segmentation.

Sources (1)

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

arXiv cs.CV Zhengyu Zou, Hao Li, Kuixuan Jiao, Liu Liu, Tingyang Xiao, Xiaolin Zhou, Fangzhou Hong, Zhizhong Su, Dingwen Zhang, Ziwei Liu 2026-07-21 arXiv:2607.19228
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-21 14:38:57.917136 UTC

TL;DR - IGGT4D is a streaming Transformer that jointly reconstructs dynamic 4D scenes and maintains object identities from continuous video. It advances online spatial intelligence by unifying camera motion, geometry, and temporally consistent instances over long sequences.

  • Processes frames sequentially using causal spatial-temporal modeling and reused historical context.
  • Incrementally tracks camera motion, scene geometry, and object identity in one representation.
  • Introduces InsScene4D-147K, covering real, synthetic, static, and dynamic scenes with geometry-guided annotations.
  • Outperforms reported streaming baselines across reconstruction, pose estimation, instance tracking, and open-vocabulary segmentation.
item →