Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
Merged summary
TL;DR - Agentic Real2Sim uses vision-language agents to turn recordings of robot-object interactions into runnable physics simulations. It aims to reduce manual real-to-simulation work and support scalable robot policy training and evaluation.
- Reconstructs geometry, object states, physical parameters, cameras, poses, and trajectories.
- Handles rigid-object manipulation, deformable-object interaction, and humanoid motion within one framework.
- Produces episodic digital twins preserving observations, interactions, and state changes.
- An open-weight VLM reportedly achieves conversion success comparable to frontier models at substantially lower cost.
Sources (1)
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
TL;DR - Agentic Real2Sim uses vision-language agents to turn recordings of robot-object interactions into runnable physics simulations. It aims to reduce manual real-to-simulation work and support scalable robot policy training and evaluation.
- Reconstructs geometry, object states, physical parameters, cameras, poses, and trajectories.
- Handles rigid-object manipulation, deformable-object interaction, and humanoid motion within one framework.
- Produces episodic digital twins preserving observations, interactions, and state changes.
- An open-weight VLM reportedly achieves conversion success comparable to frontier models at substantially lower cost.