都在问世界模型怎么落地,PixVerse把答案做成了「好玩」
TL;DR - PixVerse introduced R2, a real-time multimodal world model designed for persistent interactive entertainment rather than one-shot prompt-driven video. It combines continuous world simulation, low-latency controls, and direct model acceleration to support interactive stories, game-like environments, and digital characters.
- A unified Omni Causal AR model accepts text, speech, reference content, and keyboard controls while generating synchronized audio and video streams.
- Dynamic chunking, layered memory, noisy-history training, and an error replay bank target responsive control and long-term consistency; reported brightness drift fell from 0.201 to 0.129.
- Real-time acceleration uses direct distillation from the base model, over 90% block-sparse attention, and pyramid ultra-few-step distillation to reduce latency and high-resolution computation.
- Demonstrations include dynamically generated game events, persistent storylines after unscripted actions, and digital characters that respond to text or interrupting speech in real time.