具身新贵昆仑行斩获 WorldArena单项冠军全球亚军,世界模型全面领先
TL;DR - Kunlunx AI’s first-generation GeWu world model ranked second overall on WorldArena 2.0 Track 1 with 65.91 and first in image quality with 69.58. Its architecture explicitly models the causal chain from intent to action to consequence, aiming to improve physically consistent robot simulation and planning.
- GeWu placed in the top four on four of six evaluation dimensions among 77 participating models and scored 82.40 in interaction quality.
- Three independent Transformer towers represent intent, intervention, and consequence, with directional cross-tower attention preventing action generation from accessing future visual outcomes.
- Language is modeled autoregressively, while actions and future video frames use flow matching; noise-level adjustments support simulation, policy inference, and joint prediction in one architecture.
- Chunk-wise autoregressive generation uses recency-weighted memory sampling, absolute temporal encoding, and frame-rate scaling to maintain consistency over longer videos.