TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models
Ranking
Overall
76
Content
90
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - TempJail is a temporal jailbreak framework that makes unsafe semantics emerge across an image-to-video sequence rather than within any single frame. It exposes safety gaps in commercial video generators and shows that frame-level filtering may miss harmful meaning composed over time.
- Decomposes a malicious caption into an initial-frame visual condition and a temporal text instruction.
- Uses controlled diffusion-latent perturbations with pretrained-encoder gradient guidance to inject camouflaged visual semantics.
- Rewrites prompts into innocuous “subject-action-scene” templates that preserve temporal guidance while bypassing text safety filters.
- Across Kling, Seedance, Veo, and PixVerse, it improves attack success over prior methods by 23.3% under GPT-5.2 evaluation and 22.0% under human evaluation.
Sources (1)
TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - TempJail is a temporal jailbreak framework that makes unsafe semantics emerge across an image-to-video sequence rather than within any single frame. It exposes safety gaps in commercial video generators and shows that frame-level filtering may miss harmful meaning composed over time.
- Decomposes a malicious caption into an initial-frame visual condition and a temporal text instruction.
- Uses controlled diffusion-latent perturbations with pretrained-encoder gradient guidance to inject camouflaged visual semantics.
- Rewrites prompts into innocuous “subject-action-scene” templates that preserve temporal guidance while bypassing text safety filters.
- Across Kling, Seedance, Veo, and PixVerse, it improves attack success over prior methods by 23.3% under GPT-5.2 evaluation and 22.0% under human evaluation.