CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models
TL;DR - CertVLA is a certified defense for vision-language-action policies facing bounded physical patch and texture attacks during closed-loop control. It extends certification beyond discrete predictions to continuous, temporally correlated robot actions and, under additional correctness assumptions, can guarantee task success.
- Uses deterministic covering masks so at least one evaluated prediction remains free of any bounded-support attack.
- Certifies actions by comparing masked predictions within a calibrated region of behavioral consistency, normalized by benign variation.
- Extends per-query certificates across an entire closed-loop rollout and provides finite-sample clean-coverage guarantees.
- Its guarantees are independent of patch content, generation method, and physical transformation, with validation in simulation and real-world patch-attack experiments.