🛰️ Daily AI Frontier
‹ back to 2026-08-19

英伟达 Cosmos 3 深度拆解:它想做具身智能时代的「安卓系统」| RSS 2026

雷峰网 (AI科技评论) Multimodal & Generative 2026-08-19
Representative image for 英伟达 Cosmos 3 深度拆解:它想做具身智能时代的「安卓系统」| RSS 2026

TL;DR - Nvidia introduced Cosmos 3, an open world foundation model that unifies text, images, video, audio, and robot actions for physical AI. Its edge-optimized version aims to bring real-time perception, simulation, and control onto Jetson-class devices.

  • A single MoE architecture combines a reasoning component with a multimodal generator and supports world understanding, forward and inverse dynamics, and policy generation.
  • Training spans dynamics, inverse-dynamics, and policy modes, with synchronized positional encoding for time-aligned video, audio, and action data.
  • The family includes 4B Edge, 16B Nano, and 64B Super variants; Edge drops audio to meet on-device resource and latency constraints.
  • Nvidia reports that action conditioning improved future-visual prediction in robotics-specific domains, but does not yet claim strong global evidence across all video domains.

view merged work →