🛰️ Daily AI Frontier
‹ back to 2026-07-23

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

Research Multimodal & Generative

Ranking

Overall 55
Content 60
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - MTVDiff is a multimodal latent diffusion framework that translates thermal face images into visible-spectrum images using depth and text guidance. It improves image quality and identity preservation for face recognition under varying illumination.

  • Fuses multi-scale thermal and depth features with dual-branch cross-attention.
  • Uses gated text-to-visual alignment and spatial feature transformations to integrate semantic and multimodal priors.
  • On MCXFace and SpeakingFaces, it reports FID reductions of up to 48.3% over prior GAN- and diffusion-based methods.
  • Improves Rank-1 face-verification accuracy by up to 8.9%.

Sources (1)

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

arXiv cs.CV Zhiyuan Xia, Haojie Li, Jingyu Lin, Yiguo Qiao, Cunjian Chen 2026-07-22 arXiv:2607.19886
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-20 14:34:44.708227 UTC

TL;DR - MTVDiff is a multimodal latent diffusion framework that translates thermal face images into visible-spectrum images using depth and text guidance. It improves image quality and identity preservation for face recognition under varying illumination.

  • Fuses multi-scale thermal and depth features with dual-branch cross-attention.
  • Uses gated text-to-visual alignment and spatial feature transformations to integrate semantic and multimodal priors.
  • On MCXFace and SpeakingFaces, it reports FID reductions of up to 48.3% over prior GAN- and diffusion-based methods.
  • Improves Rank-1 face-verification accuracy by up to 8.9%.
item →