国内第一视角数据最早押注者,北大卢宗青:隐空间才是具身的路
Ranking
Overall
71
Content
80
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - BeingBeyond launched Being-H0.8, described as the first latent world-action model to jointly encode vision, touch, actions, and future state changes for robot control. Its latent-space approach targets real-time deployment at roughly 1% of pixel-based video-model training cost.
- Being-H0.8 adds tactile signals to large-scale pretraining, aiming to model physical interactions rather than merely visual observations.
- The company has curated over 500,000 hours of first-person human video, arguing it offers greater scale and diversity than robot-collected or simulated data.
- Its models predict actions and world responses directly in embedding space, avoiding costly frame generation and supporting faster inference.
- Founder Lu Zongqing cautions that embodied AI still lacks a proven paradigm comparable to next-token prediction for LLMs; even world models may not be the final answer.
Sources (1)
国内第一视角数据最早押注者,北大卢宗青:隐空间才是具身的路
Public signals
N/A
TL;DR - BeingBeyond launched Being-H0.8, described as the first latent world-action model to jointly encode vision, touch, actions, and future state changes for robot control. Its latent-space approach targets real-time deployment at roughly 1% of pixel-based video-model training cost.
- Being-H0.8 adds tactile signals to large-scale pretraining, aiming to model physical interactions rather than merely visual observations.
- The company has curated over 500,000 hours of first-person human video, arguing it offers greater scale and diversity than robot-collected or simulated data.
- Its models predict actions and world responses directly in embedding space, avoiding costly frame generation and supporting faster inference.
- Founder Lu Zongqing cautions that embodied AI still lacks a proven paradigm comparable to next-token prediction for LLMs; even world models may not be the final answer.