Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
TL;DR - DEED is a data-efficient post-training and experience-driven learning framework for deploying VLA humanoids in real retail settings. It suggests that careful systems integration can turn failed naive fine-tuning into competent operation using one GPU.
- Evaluated chip restocking with a Unitree G1-Edu humanoid and GR00T N1.6.
- Aligns control frequency, curates data, highlights task-relevant visuals, and reduces dependence on the VLA model.
- Adapts RECAP-style refinement using text-based advantage prefixes and a vision-language value function.
- Includes latent-space analysis of in- and out-of-distribution behavior.