H3-World: Turning Language Understanding into World Control
TL;DR - H3-World turns the 33B MiniMax-H3 video generator into an interactive world model controlled through temporally grounded language instructions. It achieves precise character and camera control with lightweight adaptation, suggesting pretrained video generators can efficiently acquire interactive control capabilities.
- Encodes actions as structured character and camera instructions aligned with temporal video latents.
- Uses temporal attention routing to constrain instructions to intended intervals and reduce control leakage between actions.
- Trains on 8,000 gameplay samples for 10,000 LoRA steps while updating only 0.199% of parameters.
- Preserves generation quality and generalizes to unseen scenarios without dedicated action modules.