国产具身智能创全球新纪录!以30%成本跑赢 Figure AI 45%效率,聪明的具身大脑成关键
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Chinese robotics startup 自变量机器人 (Wuwen/Bianliang) ran a live, unscripted 1-hour logistics-sorting demo where its WALL-B embodied foundation model sorted 1,816 parcels/hour at 98% accuracy — claimed ~45% above Figure AI's published 1,248/hour, using a dual-arm + standard gripper rig said to cost ~70% less than a humanoid with five-finger hands. It's a vendor-run benchmark, but it argues model capability can substitute for hardware complexity in embodied AI.
- Test conditions: randomly mixed parcels (cartons, soft mailers, cylinders, foam-packed fresh goods) in arbitrary poses on a moving conveyor, no fixed waypoints or preset trajectories, no human takeover or stoppage for the full hour.
- WALL-B is billed as the first embodied model on a "World Unified Model" (WUM) architecture — fusing vision, audio, language, touch, and action in one network rather than the modular VLA pipeline, to avoid inter-module information loss.
- Claimed behaviors: strategy switching by object properties (fast grasp for light/regular items, dual-arm cooperation for heavy boxes, side-pushing for awkward ones), separating stacked parcels before picking, flattening deformable soft bags before locating shipping labels, plus zero-shot generalization to unseen packaging.
- Positioning: the same model was previously deployed in home environments; transferring it to industrial logistics is presented as evidence for cross-scenario generalization and drop-in deployment at existing sorting stations. All figures are self-reported by the vendor and not independently verified.
Sources (1)
国产具身智能创全球新纪录!以30%成本跑赢 Figure AI 45%效率,聪明的具身大脑成关键
TL;DR - Chinese robotics startup 自变量机器人 (Wuwen/Bianliang) ran a live, unscripted 1-hour logistics-sorting demo where its WALL-B embodied foundation model sorted 1,816 parcels/hour at 98% accuracy — claimed ~45% above Figure AI's published 1,248/hour, using a dual-arm + standard gripper rig said to cost ~70% less than a humanoid with five-finger hands. It's a vendor-run benchmark, but it argues model capability can substitute for hardware complexity in embodied AI.
- Test conditions: randomly mixed parcels (cartons, soft mailers, cylinders, foam-packed fresh goods) in arbitrary poses on a moving conveyor, no fixed waypoints or preset trajectories, no human takeover or stoppage for the full hour.
- WALL-B is billed as the first embodied model on a "World Unified Model" (WUM) architecture — fusing vision, audio, language, touch, and action in one network rather than the modular VLA pipeline, to avoid inter-module information loss.
- Claimed behaviors: strategy switching by object properties (fast grasp for light/regular items, dual-arm cooperation for heavy boxes, side-pushing for awkward ones), separating stacked parcels before picking, flattening deformable soft bags before locating shipping labels, plus zero-shot generalization to unseen packaging.
- Positioning: the same model was previously deployed in home environments; transferring it to industrial logistics is presented as evidence for cross-scenario generalization and drop-in deployment at existing sorting stations. All figures are self-reported by the vendor and not independently verified.