QuantWAMs: Calibrating at the Right Granularity for World Action Models
TL;DR - QuantWAMs is a post-training quantization framework tailored to closed-loop world action models. It substantially reduces memory and accelerates targeted blocks while maintaining near-FP16 manipulation performance.
- Calibrates quantization using compatible module structure, joint video-action saliency, and reachable rollout states.
- Under W4A4-dominant quantization, simulation means differ from FP16 by only 0.2–0.7 percentage points.
- Cuts peak weight-and-activation memory for targeted blocks to about 29% of FP16.
- Delivers 1.4–1.6× block-level speedups and demonstrates feasibility on three real-robot tasks.