LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
Ranking
Overall
70
Content
80
Popularity
45
Observed public metrics from 1 member.
Merged summary
TL;DR - LaP-Forensics detects and localizes synthetic-image artifacts by combining RGB content with residuals from Stable Diffusion inversion and reconstruction. It improves cross-generator forensic reasoning while acknowledging unresolved textual-faithfulness and post-processing robustness issues.
- A structured Where-What-Why model produces textual analysis and artifact masks from separately encoded RGB and residual features.
- Training combines supervised fine-tuning with GRPO rewards for mask overlap, output structure, and references to residual evidence.
- A separate image-level head fuses RGB and residual features for classification.
- Experiments report cross-generator detection on UniversalFakeDetect and competitive localization on SynthScars.
Sources (1)
LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - LaP-Forensics detects and localizes synthetic-image artifacts by combining RGB content with residuals from Stable Diffusion inversion and reconstruction. It improves cross-generator forensic reasoning while acknowledging unresolved textual-faithfulness and post-processing robustness issues.
- A structured Where-What-Why model produces textual analysis and artifact masks from separately encoded RGB and residual features.
- Training combines supervised fine-tuning with GRPO rewards for mask overlap, output structure, and references to residual evidence.
- A separate image-level head fuses RGB and residual features for classification.
- Experiments report cross-generator detection on UniversalFakeDetect and competitive localization on SynthScars.