🛰️ Daily AI Frontier
‹ back to 2026-07-23

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

arXiv cs.CV Multimodal & Generative Zhiyuan Xia, Haojie Li, Jingyu Lin, Yiguo Qiao, Cunjian Chen 2026-07-22

TL;DR - MTVDiff is a multimodal latent diffusion framework that translates thermal face images into visible-spectrum images using depth and text guidance. It improves image quality and identity preservation for face recognition under varying illumination.

  • Fuses multi-scale thermal and depth features with dual-branch cross-attention.
  • Uses gated text-to-visual alignment and spatial feature transformations to integrate semantic and multimodal priors.
  • On MCXFace and SpeakingFaces, it reports FID reductions of up to 48.3% over prior GAN- and diffusion-based methods.
  • Improves Rank-1 face-verification accuracy by up to 8.9%.

view merged work →