Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models
Merged summary
TL;DR - This paper introduces Riemannian flow matching in VGGT’s latent space to generate plausible 3D scenes from sparse, unposed images. It combines a geometric foundation model’s learned priors with generative modeling while avoiding a fixed output representation.
- Models VGGT tokens on a product manifold of four high-dimensional hyperspheres rather than using unsuitable Euclidean flow matching.
- Keeps generated latent tokens valid for VGGT’s frozen decoding heads and multi-scale encoder.
- Avoids committing to representations such as Gaussian splats, meshes, or video-VAE latents.
- Reports strong appearance and aggregated 3D geometry performance on RealEstate10K, ScanNet++, and ETH3D against recent baselines.
Sources (1)
Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models
TL;DR - This paper introduces Riemannian flow matching in VGGT’s latent space to generate plausible 3D scenes from sparse, unposed images. It combines a geometric foundation model’s learned priors with generative modeling while avoiding a fixed output representation.
- Models VGGT tokens on a product manifold of four high-dimensional hyperspheres rather than using unsuitable Euclidean flow matching.
- Keeps generated latent tokens valid for VGGT’s frozen decoding heads and multi-scale encoder.
- Avoids committing to representations such as Gaussian splats, meshes, or video-VAE latents.
- Reports strong appearance and aggregated 3D geometry performance on RealEstate10K, ScanNet++, and ETH3D against recent baselines.