GraphVid: Interactive Graph-Controllable Video Generation
Merged summary
TL;DR - GraphVid is a graph-conditioned image-to-video model that lets users control complex multi-object interactions through structured interaction graphs. It offers stronger controllability and video quality while using less training data and fewer trainable parameters than prior motion-control methods.
- Replaces cumbersome object-trajectory drawing with semantic graph-based interaction control.
- Introduces GraphVid-Bench, a large interaction-focused video dataset with structured relational annotations.
- Versus Motion-I2V, reduces FID by up to 39.9% and FVD by 37.6%.
- Improves PSNR from 9.87 to 15.98 and SSIM from 0.38 to 0.61.
Sources (1)
GraphVid: Interactive Graph-Controllable Video Generation
TL;DR - GraphVid is a graph-conditioned image-to-video model that lets users control complex multi-object interactions through structured interaction graphs. It offers stronger controllability and video quality while using less training data and fewer trainable parameters than prior motion-control methods.
- Replaces cumbersome object-trajectory drawing with semantic graph-based interaction control.
- Introduces GraphVid-Bench, a large interaction-focused video dataset with structured relational annotations.
- Versus Motion-I2V, reduces FID by up to 39.9% and FVD by 37.6%.
- Improves PSNR from 9.87 to 15.98 and SSIM from 0.38 to 0.61.