🛰️ Daily AI Frontier
‹ back to 2026-09-10

浙大 × 字节 × 港中深最新研究:告别高斯堆砌,终结 3D 冗余|ECCV 2026

Research Multimodal & Generative

Ranking

Overall 75
Content 85
Popularity 50

Observed public metrics from 1 member.

Representative image for 浙大 × 字节 × 港中深最新研究:告别高斯堆砌,终结 3D 冗余|ECCV 2026

Merged summary

TL;DR - PointSplat, an ECCV 2026 paper from Zhejiang University, ByteDance, and CUHK-Shenzhen, proposes a compact, feed-forward 3D Gaussian Splatting framework for reconstructing humans from sparse multi-view images. It reduces redundant representations by organizing Gaussians around the person rather than independently around each camera view.

  • Replaces view-centric prediction, which repeatedly encodes the same subject across views, with a shared human-centric 3D representation.
  • Estimates a coarse human proxy and uses ray casting to remove irrelevant or invisible points before Gaussian prediction.
  • A Point-Image Transformer fuses multi-view geometry and appearance into Gaussian position, scale, rotation, opacity, and visual attributes.
  • Targets compact, generalizable 3D human reconstruction for bandwidth- and compute-sensitive applications such as immersive live streaming.

Sources (1)

浙大 × 字节 × 港中深最新研究:告别高斯堆砌,终结 3D 冗余|ECCV 2026

雷峰网 (AI科技评论) 2026-09-10 arXiv:2606.32036
Public signals Hugging Face upvotes 1
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:21:25.480700 UTC

TL;DR - PointSplat, an ECCV 2026 paper from Zhejiang University, ByteDance, and CUHK-Shenzhen, proposes a compact, feed-forward 3D Gaussian Splatting framework for reconstructing humans from sparse multi-view images. It reduces redundant representations by organizing Gaussians around the person rather than independently around each camera view.

  • Replaces view-centric prediction, which repeatedly encodes the same subject across views, with a shared human-centric 3D representation.
  • Estimates a coarse human proxy and uses ray casting to remove irrelevant or invisible points before Gaussian prediction.
  • A Point-Image Transformer fuses multi-view geometry and appearance into Gaussian position, scale, rotation, opacity, and visual attributes.
  • Targets compact, generalizable 3D human reconstruction for bandwidth- and compute-sensitive applications such as immersive live streaming.
item →