🛰️ Daily AI Frontier
‹ back to 2026-09-10

全球首个可仿真的人–场景交互重建框架 HSImul3R:让人类视频真正成为机器人技能来源

Research Embodied Robotics

Ranking

Overall 79
Content 85
Popularity 65

Observed public metrics from 1 member.

Representative image for 全球首个可仿真的人–场景交互重建框架 HSImul3R:让人类视频真正成为机器人技能来源

Merged summary

TL;DR - HSImul3R is an ECCV 2026-accepted framework that reconstructs simulation-ready human–scene interactions from sparse or monocular video, using physics feedback to optimize both motion and scene geometry. It could turn human videos into physically valid skills transferable to humanoid robots.

  • Its physics-in-the-loop design combines scene-targeted reinforcement learning for interaction-aware human motion with direct simulation reward optimization for 3D scenes.
  • On HSIBench, interaction stability reached 53.68%, 30.56%, and 13.92% across Easy, Medium, and Hard tasks, versus 10.52%, 4.50%, and 2.66% for HSfM.
  • The framework reduced human–scene 3D interpenetration from 69.51% to 22.90%.
  • Optimized motions were retargeted to a Unitree G1 and deployed through a simulation-to-real whole-body control pipeline.

Sources (1)

全球首个可仿真的人–场景交互重建框架 HSImul3R:让人类视频真正成为机器人技能来源

量子位 量子位的朋友们 2026-09-10 arXiv:2603.15612
Public signals Hugging Face upvotes 51
Providers: Hugging Face · Upvotes 51 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:21:28.705961 UTC

TL;DR - HSImul3R is an ECCV 2026-accepted framework that reconstructs simulation-ready human–scene interactions from sparse or monocular video, using physics feedback to optimize both motion and scene geometry. It could turn human videos into physically valid skills transferable to humanoid robots.

  • Its physics-in-the-loop design combines scene-targeted reinforcement learning for interaction-aware human motion with direct simulation reward optimization for 3D scenes.
  • On HSIBench, interaction stability reached 53.68%, 30.56%, and 13.92% across Easy, Medium, and Hard tasks, versus 10.52%, 4.50%, and 2.66% for HSfM.
  • The framework reduced human–scene 3D interpenetration from 69.51% to 22.90%.
  • Optimized motions were retargeted to a Unitree G1 and deployed through a simulation-to-real whole-body control pipeline.
item →