自进化WAM来了!清华AIR联手域变换提出具身In-Context Causal Learning
TL;DR - Tsinghua AIR and startup Yubianhuan introduced Zeva, an embodied world-action model that learns from action outcomes and human demonstrations at deployment time without updating its weights. Its causal memory raised cumulative success from 26% to 73% across repeated attempts on a simulated benchmark and improved real-world robotic lab tasks.
- Zeva encodes visual state, actions, and resulting state changes into causal interaction signals that capture action-effect relationships.
- Dual-timescale memory maintains short-term execution context while retaining reusable evidence from failures, corrections, and successes across attempts.
- Retrieved evidence is injected as a causal prompt into the frozen Cosmos3 action generator, enabling zero-gradient adaptation and one-shot learning from human demonstrations.
- On ChemLab-Evo, repeated attempts improved success from 65% to 100% for picking up test tubes, 25% to 70% for placing beakers, and 30% to 80% for pouring water.