🛰️ Daily AI Frontier
‹ back to 2026-07-29

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

Research Deepfake Detection

Ranking

Overall 70
Content 80
Popularity 45

Observed public metrics from 1 member.

Representative image for LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

Merged summary

TL;DR - LaP-Forensics detects and localizes synthetic-image artifacts by combining RGB content with residuals from Stable Diffusion inversion and reconstruction. It improves cross-generator forensic reasoning while acknowledging unresolved textual-faithfulness and post-processing robustness issues.

  • A structured Where-What-Why model produces textual analysis and artifact masks from separately encoded RGB and residual features.
  • Training combines supervised fine-tuning with GRPO rewards for mask overlap, output structure, and references to residual evidence.
  • A separate image-level head fuses RGB and residual features for classification.
  • Experiments report cross-generator detection on UniversalFakeDetect and competitive localization on SynthScars.

Sources (1)

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

arXiv cs.CV Can Wang, Yuhao Wang, Yushe Cao, Canran Xiao, Fei Shen 2026-07-28 arXiv:2607.25962
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-18 14:31:58.210639 UTC

TL;DR - LaP-Forensics detects and localizes synthetic-image artifacts by combining RGB content with residuals from Stable Diffusion inversion and reconstruction. It improves cross-generator forensic reasoning while acknowledging unresolved textual-faithfulness and post-processing robustness issues.

  • A structured Where-What-Why model produces textual analysis and artifact masks from separately encoded RGB and residual features.
  • Training combines supervised fine-tuning with GRPO rewards for mask overlap, output structure, and references to residual evidence.
  • A separate image-level head fuses RGB and residual features for classification.
  • Experiments report cross-generator detection on UniversalFakeDetect and competitive localization on SynthScars.
item →