🛰️ Daily AI Frontier
‹ back to 2026-08-23

上海 AI Lab 提出 OccAnyScene:一个模型统一室内外 3D 语义占据预测

Research 3D Scene Understanding

Ranking

Overall 68
Content 80
Popularity 40

Observed public metrics from 1 member.

Representative image for 上海 AI Lab 提出 OccAnyScene:一个模型统一室内外 3D 语义占据预测

Merged summary

TL;DR - OccAnyScene is a unified 3D semantic occupancy model that handles indoor rooms and outdoor roads despite differing cameras, spatial scales, voxel grids, and label systems. Joint training nearly matches scene-specific models, suggesting a path toward general-purpose occupancy foundation models.

  • Uses pixel frustums as scale-adaptive geometric units and decodes them into 3D Gaussians representing visible and occluded regions.
  • Shares all model parameters across indoor and outdoor scenes except dataset-specific semantic mapping matrices.
  • The jointly trained model reaches 59.51% mIoU on Occ-ScanNet and 22.87% on SurroundOcc-nuScenes, within 0.41 and 0.19 points of scene-specific versions.
  • Evaluation covers only two benchmarks, and areas outside camera-frustum coverage still require supplementary spatial queries.

Sources (1)

上海 AI Lab 提出 OccAnyScene:一个模型统一室内外 3D 语义占据预测

WeChat: 自动驾驶之心 2026-08-20 arXiv:2608.08696
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-25 03:06:56.733442 UTC

TL;DR - OccAnyScene is a unified 3D semantic occupancy model that handles indoor rooms and outdoor roads despite differing cameras, spatial scales, voxel grids, and label systems. Joint training nearly matches scene-specific models, suggesting a path toward general-purpose occupancy foundation models.

  • Uses pixel frustums as scale-adaptive geometric units and decodes them into 3D Gaussians representing visible and occluded regions.
  • Shares all model parameters across indoor and outdoor scenes except dataset-specific semantic mapping matrices.
  • The jointly trained model reaches 59.51% mIoU on Occ-ScanNet and 22.87% on SurroundOcc-nuScenes, within 0.41 and 0.19 points of scene-specific versions.
  • Evaluation covers only two benchmarks, and areas outside camera-frustum coverage still require supplementary spatial queries.
item →