🛰️ Daily AI Frontier
‹ back to 2026-08-30

TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

Research Multimodal & Generative

Ranking

Overall 76
Content 90
Popularity 42

Observed public metrics from 1 member.

Representative image for TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

Merged summary

TL;DR - TempJail is a temporal jailbreak framework that makes unsafe semantics emerge across an image-to-video sequence rather than within any single frame. It exposes safety gaps in commercial video generators and shows that frame-level filtering may miss harmful meaning composed over time.

  • Decomposes a malicious caption into an initial-frame visual condition and a temporal text instruction.
  • Uses controlled diffusion-latent perturbations with pretrained-encoder gradient guidance to inject camouflaged visual semantics.
  • Rewrites prompts into innocuous “subject-action-scene” templates that preserve temporal guidance while bypassing text safety filters.
  • Across Kling, Seedance, Veo, and PixVerse, it improves attack success over prior methods by 23.3% under GPT-5.2 evaluation and 22.0% under human evaluation.

Sources (1)

TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

arXiv cs.CV Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang 2026-08-27 arXiv:2608.26971
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-12 14:18:29.890092 UTC

TL;DR - TempJail is a temporal jailbreak framework that makes unsafe semantics emerge across an image-to-video sequence rather than within any single frame. It exposes safety gaps in commercial video generators and shows that frame-level filtering may miss harmful meaning composed over time.

  • Decomposes a malicious caption into an initial-frame visual condition and a temporal text instruction.
  • Uses controlled diffusion-latent perturbations with pretrained-encoder gradient guidance to inject camouflaged visual semantics.
  • Rewrites prompts into innocuous “subject-action-scene” templates that preserve temporal guidance while bypassing text safety filters.
  • Across Kling, Seedance, Veo, and PixVerse, it improves attack success over prior methods by 23.3% under GPT-5.2 evaluation and 22.0% under human evaluation.
item →